How much CO2 does AI produce?
There is no universal grams-per-prompt figure. AI emissions depend on electricity, grid carbon intensity and whether training and hardware are counted.

There is no trustworthy universal amount of CO2 for one AI prompt or for AI worldwide. For a single response, the basic calculation is electricity used multiplied by the carbon intensity of the electricity supply. Both inputs vary. A short answer from a small model and a long, tool-using response from a frontier system are different jobs, and the same job can have different emissions on different grids.
At global scale, public evidence is usually for all data centres rather than AI alone. The International Energy Agency projects electricity-related data-centre emissions reaching around 320 million tonnes of CO2 in 2030 in its base case. AI is a major growth driver, but storage, cloud computing and other digital services are included, so 320 million tonnes is not an AI total.
Turning watt-hours into grams of CO2
The arithmetic is straightforward once the assumptions are chosen. Convert watt-hours to kilowatt-hours, then multiply by grams of CO2-equivalent per kilowatt-hour. If an ordinary text query uses roughly 0.3 Wh, as Epoch AI estimates for a typical GPT-4o request, its operational emissions would be 0.0003 kWh multiplied by the relevant grid factor.
There is no universal factor to finish that equation. A renewable- and nuclear-heavy grid, a gas-heavy grid and a coal-heavy grid produce different answers. The factor can also change hour by hour.
Any published gram figure should therefore show both the energy estimate and carbon-intensity source rather than hiding them behind a decimal.
A message can be much larger than average
The roughly 0.3 Wh estimate describes a typical text request under Epoch’s assumptions. Long context changes the computation. So do extended reasoning, multiple model calls, web tools and media generation. The model’s architecture, hardware and utilisation also influence energy per token.
This is why an answer such as “AI produces X grams per prompt” is incomplete even before carbon is considered. The prompt is an interface event, not a standard workload. A service average may be useful for rough planning, but it cannot serve as a meter reading for every request.
The energy cost of one AI message explains how input, output and hidden backend steps widen the range.
Training does not fit neatly into a prompt
Training creates or updates a model through a large, finite computing job. Inference is the continuing work when the model is used. A per-response estimate normally covers inference electricity, perhaps with data-centre overhead, and leaves training outside the boundary.
Training emissions can be allocated across messages, but the divisor is unknown until the model retires. A model used billions of times would receive a much smaller training allocation per message than one trained and quickly abandoned. Development experiments make the numerator uncertain too.
For that reason, training and inference emissions are clearer when reported separately. Combining them can be appropriate for a retrospective life-cycle study, but not when an uncertain allocation is presented as direct per-message measurement.
Operational electricity is only one carbon source. Manufacturing accelerators and memory, building servers, producing concrete and steel, constructing data centres and moving equipment through supply chains all create embodied emissions. Backup generation and equipment replacement may add further impacts.
These sources tend to appear in corporate Scope 3 inventories rather than per-model figures. Microsoft’s 2025 sustainability report reported that its total emissions remained above its 2020 baseline amid AI and cloud expansion, even as it described progress on direct operational emissions and carbon-free electricity procurement.
That does not reveal the CO2 from one product. It demonstrates why a clean operational electricity claim is not a complete AI footprint.
Carbon accounting can produce different answers
A location-based calculation follows the emissions intensity of the grid where electricity is consumed. A market-based inventory can recognise contracts, certificates and other instruments associated with cleaner generation. Hourly matching asks a stricter question than annual matching: was carbon-free electricity available at the time and place of use?
The IEA states that its data-centre supply analysis uses the physical electricity mix rather than companies’ contractual mix. A corporate report may legitimately use recognised market-based accounting and arrive at a lower operational result. When numbers conflict, inspect the method before deciding one must be false.
Our AI carbon footprint explainer covers these boundaries in more detail.
What the global forecasts do and do not say
The IEA’s base case puts all data-centre electricity use at about 415 TWh in 2024 and about 945 TWh in 2030. It expects AI to be the most important driver of growth alongside other digital services. Its electricity-related emissions rise more slowly than energy demand as lower-emissions generation expands, although fossil fuels still supply a significant share of near-term growth.
Those are scenarios, not measurements of the future. Efficiency, AI uptake, construction and the generation mix could change the path. The IEA publishes alternative cases for that reason.
Nor should a global share minimise local effects. Data-centre demand is concentrated, so a new facility can materially affect a regional grid even when the sector remains a modest part of global emissions.
A defensible way to report one response
State the model and workload, estimate or measure the inference electricity, say whether facility overhead is included, then disclose the grid factor and time basis. Keep training and embodied emissions visible as exclusions unless there is evidence to allocate them. Round to match the uncertainty.
Treechat follows that narrower operational approach. It estimates electricity from the model and token count, adds data-centre overhead and applies a stated average grid carbon intensity. Providers do not supply a per-response power meter or exact facility, so the result remains an estimate. The current constants and limitations are at /methodology.
The answer to “how much CO2 does AI produce?” is therefore conditional. For one response, show the equation and assumptions. For the world, use data-centre scenarios without relabelling them as AI-only totals. Precision without those boundaries is decoration, not knowledge.