Treechat
Menu
Essayer gratuitement
← All notes

How much energy does a ChatGPT query use?

The honest answer is a range, not a single number. Here is what is public, what is estimated and what gets left out.

By Treechat6 min read
A single prompt ribbon passes through a processor and a physical energy meter.

The short answer is that a normal text query probably uses a fraction of a watt-hour, but there is no universal “ChatGPT query”. A request to fix a sentence and a deep-research task that reads dozens of sources are different computing jobs. Treating them as interchangeable makes any precise-looking figure less useful than it appears.

That does not mean we know nothing. In 2025, OpenAI chief executive Sam Altman wrote that an average ChatGPT query uses about 0.34 watt-hours. Around the same time, independent researchers at Epoch AI estimated roughly 0.3 watt-hours for a typical GPT-4o query. The agreement is notable, but it is not the same as a public, independently audited measurement. Altman did not publish the calculation behind his figure, while Epoch had to make assumptions about the model, hardware, utilisation and token count.

So 0.3–0.34 Wh is a useful reference point for an ordinary text exchange of the kind those estimates describe. It is not a meter reading for your particular prompt, and it should not be stretched to cover every mode or model.

Why one query can be very different from another

An AI response is generated token by token. Broadly, more input to process and more output to generate mean more work. A short factual answer may finish quickly. A long conversation carries previous context forward. Uploading a substantial document adds an upfront processing cost. A reasoning model may generate hidden working tokens before it shows an answer. Search, image generation and agentic tools can call several systems rather than one.

Epoch’s worked examples make the size of this variation clear. Its model put a query with a 10,000-token input at around 2.5 Wh and one with a 100,000-token input near 40 Wh, under its stated assumptions. Those are not estimates for the average chat. They show why the word “query” is too coarse on its own.

The model matters as well. A smaller model generally performs fewer calculations per generated token than a large one, although architecture, quantisation, hardware and software optimisation complicate the comparison. Mixture-of-experts models activate only part of their total parameter count for each token. Modern accelerators can do more work per unit of electricity than older hardware. Efficient batching lets one server handle several users together, but an empty or lightly used server may be less efficient per answer.

Then there is the data centre around the chips. Cooling, networking and power conversion consume electricity too. Researchers commonly account for that overhead with a power usage effectiveness factor. The result also depends on where and when the electricity is generated: a watt-hour has the same energy everywhere, but its associated emissions vary with the grid mix.

What is actually being counted?

Most per-query figures describe operational electricity for inference: the energy used while the service produces an answer, plus some allowance for data-centre overhead. They may not include the user’s phone or laptop, network transmission, storage, model training, manufacture of the chips and buildings, or the infrastructure kept ready between requests.

There is no single correct boundary for every question. If you are comparing two inference options, operational energy per response is useful. If you are preparing a company-wide carbon inventory, a wider life-cycle boundary may be more appropriate. Problems begin when a narrow estimate is presented as the whole environmental cost—or when a broad estimate is compared with a narrow one.

Training deserves separate treatment. It is a large, occasional job whose impact can be allocated across an unknown number of future uses. Divide it by more messages and its per-message share shrinks; include a model that is quickly replaced and it rises. An allocation can help with accounting, but it is not electricity drawn at the moment you press send.

The same caution applies to water and carbon. Water use depends on cooling technology and the electricity supply chain. Carbon depends on location, timing and whether the accounting uses the physical grid mix or contractual renewable-energy purchases. Converting Wh into grams of carbon without stating those choices hides important uncertainty.

How the old 3 Wh figure travelled

For several years, articles often said a ChatGPT request used about 3 Wh, sometimes paired with the claim that this was ten times a Google search. The figure came from an early estimate using older A100 hardware, a GPT-3.5-scale model and an assumed 2,000 output tokens. Epoch AI explains why it considers those assumptions unrepresentative of typical modern use in its methodology and comparison.

Updating an estimate when hardware and services change is good practice. It does not prove that AI energy use is irrelevant. It means a popular per-query number aged badly.

The comparison with web search is shaky for another reason: the commonly quoted search figure came from Google in 2009. Search now involves different infrastructure and sometimes includes generated summaries. Putting a recent AI estimate beside an old search estimate creates an apparently simple ratio from unlike measurements.

How numbers get misused in both directions

One misuse is to take a high-end or outdated estimate, multiply it by a huge number of imagined requests, and present the result as a measured total. That can exaggerate ordinary text use and distract from the workloads that genuinely are more intensive.

The opposite misuse starts with a low average and concludes there is nothing to consider. A small amount multiplied across large and growing use is still system demand. The International Energy Agency’s work on energy and AI looks at data-centre demand at that system level. Per-query efficiency and total electricity growth can both be true when usage expands and new tasks require more compute.

Another mistake is false precision. “0.34 Wh” sounds measured to two decimal places, but the public statement describes an average across an undisclosed mix of requests. It is more honest to retain the context and uncertainty than to turn it into a universal constant.

How Treechat estimates a message

Treechat cannot read a power meter attached to every response. Providers do not expose that measurement. We estimate instead, and label the result as an estimate.

Our calculation starts with the model used and the tokens processed and generated. It applies a public model-based energy methodology, adds data-centre overhead, and converts electricity into estimated emissions using an average grid carbon intensity. The assumptions and current constants are set out on our methodology page.

We route everyday requests to a smaller model by default because model choice is one lever we can actually pull. We also show an estimated receipt on the answer rather than claiming an exact measurement. Where evidence is incomplete, we round against ourselves. We would rather under-claim a saving than make a neat promise the underlying data cannot support.

The most defensible answer, then, is not “a ChatGPT query uses exactly X”. It is: ordinary text queries appear to sit around a fraction of a watt-hour under recent public estimates, while long-context, reasoning and tool-heavy work can use much more. Ask which model, how many tokens, what hardware and what boundary. Without those details, the number is a signpost, not a fact about every query.

How much energy does a ChatGPT query use? · Treechat blog