The Energy Cost of an AI Query: Vendor Numbers Versus Independent Measurement
OpenAI and Google now publish per-prompt energy figures, but the measurement boundaries differ, the underlying estimates still vary by an order of magnitude, and aggregate data-centre demand keeps climbing regardless.
Ask how much energy a single AI chatbot reply costs and you will get an answer with real confidence attached to it β 0.3 watt-hours, 0.34 watt-hours, 3 watt-hours, "less than a Google search," "ten times a Google search." The confidence is misplaced. Two of the companies that actually operate these systems have now published their own per-query figures, and independent researchers have built parallel measurement frameworks from the outside. The numbers do not agree, the measurement boundaries are rarely the same, and the aggregate trend β electricity drawn by AI-focused data centres β is moving in a direction that per-query efficiency gains have not offset.
In June 2025, OpenAI CEO Sam Altman wrote in a blog post that "the average query uses about 0.34 watt-hours, about what an oven would use in a little over one second, or a high-efficiency lightbulb would use in a couple of minutes," adding a water figure of 0.000085 gallons per query, as stated in his post "The Gentle Singularity". It was the first figure OpenAI had put its name to. It was also, notably, not a methodology: there is no description of what counts as an "average" query, which model served it, whether the figure includes idle-capacity overhead or cooling, or whether the more compute-intensive reasoning and "deep research" modes are folded in or excluded.
Google went further two months later. An August 2025 technical paper, "Measuring the environmental impact of delivering AI at Google Scale," reported that the median Gemini Apps text prompt used 0.24 watt-hours of electricity, 0.03 grams of CO2-equivalent, and 0.26 millilitres of water, based on May 2025 production data. Unlike Altman's figure, Google's paper specifies a measurement boundary: active AI accelerator power, host system energy, idle machine capacity provisioned for reliability, and data-centre overhead (cooling and power delivery), combined into a full-stack accounting rather than just the chip doing the matrix multiplication. Google also reported that this same per-prompt footprint had fallen 33-fold in the twelve months before the measurement, attributing the drop to model and serving efficiency work β a claim about a company's own infrastructure that has not been independently replicated.

What vendors report is not what independent researchers can check
The gap between OpenAI's and Google's numbers is real but modest β 0.34 Wh versus 0.24 Wh, both for closed, proprietary models running on unspecified hardware at unspecified utilisation. Neither figure is independently verifiable, because neither company discloses model size, token counts, GPU allocation, or the batching and caching arrangements that determine actual energy draw per request. A single-vendor number, however precisely stated, is a claim, not a measurement anyone outside the company can reproduce.
Independent estimates, built without access to production infrastructure, span a much wider range. Epoch AI's February 2025 analysis arrived at roughly 0.3 watt-hours for a typical GPT-4o query, assuming around 100 billion active parameters, an output of roughly 500 tokens, and NVIDIA H100 GPUs run at an estimated 10% compute utilisation. That figure was a substantial revision downward from the estimate that had circulated for the previous two years: Alex de Vries's 2023 paper in Joule had put a ChatGPT query at roughly 3 watt-hours, based on higher assumed output length and older A100 hardware. Both are estimates built on public assumptions about model architecture and hardware, not on measured production traffic β which is exactly the gap Google's paper was designed to close for its own systems, and exactly what remains open for OpenAI's.
| Source | Type | Year | Figure (typical text query) | Basis |
|---|---|---|---|---|
| de Vries, Joule | Independent estimate | 2023 | ~3 Wh | Assumed A100 hardware, ~2,000-token output |
| Epoch AI | Independent estimate | 2025 | ~0.3 Wh | Assumed H100 hardware, ~500-token output, 10% utilisation |
| OpenAI (Sam Altman) | Vendor disclosure | 2025 | 0.34 Wh | Undisclosed methodology and boundary |
| Vendor disclosure | 2025 | 0.24 Wh | Full-stack: accelerator, host, idle capacity, data-centre overhead |
The real variance is by task, not by chatbot
Comparing single numbers for "a query" obscures a larger and better-documented source of variation: what kind of task the model is doing. Luccioni, Jernite and Strubell's 2024 study at the ACM Conference on Fairness, Accountability, and Transparency measured energy use across ten task types on shared infrastructure and found roughly 0.002 kilowatt-hours per 1,000 inferences for text classification, 0.047 kWh per 1,000 for text generation, and 2.9 kWh per 1,000 for image generation β a difference of roughly three orders of magnitude between the cheapest and most expensive task type, all on the same hardware. The same paper found that using a general-purpose generative model for a task a small, purpose-built discriminative model can already do β such as classifying text β multiplied emissions roughly tenfold to thirtyfold for identical accuracy, because the generative model has to run a much larger network to produce the same yes/no answer.
A follow-up benchmarking effort, Jegham and colleagues' 2025 comparison of GPT-4o, Claude, Llama and DeepSeek variants, extended this task-level approach to current frontier models across energy, water and carbon simultaneously, and again found that response length and reasoning depth β not the brand of chatbot β were the dominant drivers of per-query cost. This matters because most public figures, including both vendor disclosures above, describe a "typical" or "median" text prompt. None of the public figures for ChatGPT or Gemini quote a number for a reasoning-heavy request, an agentic multi-step task, or an image or video generation call, even though those are the categories every source agrees cost the most.
Why the aggregate trend does not match the per-query trend
Per-query efficiency has been improving quickly. The International Energy Agency's 2025 "Energy and AI" analysis states that software and hardware advances have driven energy use per AI task down by at least an order of magnitude annually in recent years, to the point that a simple text query now typically uses less electricity than running a television for the same span of time. That is consistent with the direction, if not the exact magnitude, of Epoch AI's and Google's downward revisions.

But total demand moves the other way. Electricity consumption from data centres rose 17% globally in 2025, with consumption from AI-focused data centres specifically climbing 50% in the same year, according to the IEA's April 2026 report "Key Questions on Energy and AI," which reviewed full-year 2025 outcomes β both figures well above the roughly 3% growth rate of global electricity demand overall. The same report projects data-centre electricity use to roughly double from 485 terawatt-hours in 2025 to 950 terawatt-hours by 2030, with the AI-focused share of that tripling. The agency's framing is that AI energy demand is the net result of three trends moving simultaneously: efficiency per task improving quickly, adoption volume growing faster than efficiency gains can offset it, and newer capabilities β video generation, extended reasoning, agentic multi-step tasks β that the IEA states can consume "hundreds or thousands of times more energy per query than simple text generation." A falling per-query number and a rising total are not a contradiction; they describe the same system from two different angles, and reporting only the falling number is incomplete.
Water and carbon are not the same accounting problem as electricity
Energy figures dominate public discussion, but water consumption follows a different logic tied to cooling method and grid location rather than compute alone. The foundational methodology here is Li, Yang, Islam and Ren's 2023 paper, which estimated that training GPT-3 in a modern US data centre could directly evaporate on the order of 700,000 litres of freshwater, and that a conversation of 10 to 50 exchanges with a chatbot of that era corresponded to roughly 500 millilitres of water consumption through cooling. Google's and OpenAI's 2025 disclosures both include water figures (0.26 mL and 0.000085 gallons, or about 0.32 mL, per query respectively) that are far below this earlier per-conversation estimate, consistent with the same efficiency-gain story told for energy β though again, none of these figures are directly comparable across companies because water intensity depends heavily on which specific data centre, cooling technology and local grid served the request, none of which are disclosed at the level this methodology requires to check.
What can and cannot be said with the current evidence
Three things hold up against the sources above. First, per-query energy and water use for ordinary text chatbot interactions has genuinely fallen since the earliest 2023 estimates, by roughly an order of magnitude on both the vendor-reported and independently-modelled numbers β this is the one point where OpenAI, Google, Epoch AI and the IEA all point the same direction. Second, the remaining spread between sources, even after that improvement, is still three-to-tenfold depending on assumptions about hardware, utilisation and output length, and neither OpenAI's nor Google's disclosure includes enough methodological detail for an outside party to fully reproduce or audit it β Google's paper is the more transparent of the two, specifying its measurement boundary, but it is still self-reported and unaudited. Third, none of the per-query figures in circulation β vendor or independent β cover the workloads that every source agrees cost the most: extended reasoning, agentic execution, and image or video generation. A researcher who needs an energy figure for a specific class of AI task should treat "0.3 watt-hours" as a number that applies, at best, to a short text reply from a particular model in a particular year, and should look for a task-matched, methodology-disclosed source before using it for anything else.
- Sam Altman. The Gentle Singularity. Sam Altman's blog, 2025. link
- Jared Kaplan et al. (Epoch AI staff). How much energy does ChatGPT use?. Epoch AI, 2025. link
- Alex de Vries. The growing energy footprint of artificial intelligence. Joule, 2023. doi:10.1016/j.joule.2023.09.004
- Alexandra Sasha Luccioni, Yacine Jernite, Emma Strubell. Power Hungry Processing: Watts Driving the Cost of AI Deployment?. ACM Conference on Fairness, Accountability, and Transparency (FAccT '24), 2024. doi:10.1145/3630106.3658542
- Google (Jeff Dean et al.). Measuring the environmental impact of delivering AI at Google Scale. arXiv / Google, 2025. link
- International Energy Agency. Energy and AI. IEA, 2025. link
- International Energy Agency. Key Questions on Energy and AI. IEA, 2026. link
- Pengfei Li, Jianyi Yang, Mohammad A. Islam, Shaolei Ren. Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models. arXiv, 2023. link
- Nidhal Jegham, Marwan Abdelatti, Chan Young Koh, Lassad Elmoubarki, Abdeltawab Hendawi. How Hungry is AI? Benchmarking Energy, Water, and Carbon Footprint of LLM Inference. arXiv, 2025. link