AI water usage per query: what can we actually know?
AI water estimates range from 0.26 ml to tens of millilitres per text query. The gap comes from systems, locations and accounting boundaries.

There is no single water-use figure for an AI query. Recent operator disclosures put ordinary text prompts at about 0.26 ml for the median Gemini Apps prompt and 0.32 ml for an average ChatGPT query. A 2023 academic model produced a much higher range of roughly 10–25 ml per language-model response under particular locations and system boundaries.
These estimates should not be averaged or treated as measurements of the same thing. They cover different models, hardware, cooling, grids, dates and definitions of water. A query’s length and tools matter too. The most accurate consumer answer is a sourced range with its boundary attached.
The two low operator figures
Google reported 0.26 ml of water for a median Gemini Apps text-generation prompt in May 2025. Its accompanying analysis also reported 0.24 Wh of energy and described a comprehensive serving boundary including supporting infrastructure.
OpenAI chief executive Sam Altman stated that an average ChatGPT query used 0.000085 US gallons, approximately 0.32 ml. The post did not publish a calculation or define the underlying request mix.
The figures are close, but that does not establish a universal rate. One is a measured point-in-time median for a named app; the other is an unexplained fleet average. Median and mean are different statistics, and both services continue to change.
The higher research estimate
In Making AI Less “Thirsty”, researchers modelled direct cooling and indirect electricity-related water for large language models. Their widely cited example assigned a 500 ml bottle to about 20–50 medium-length responses, equivalent to around 10–25 ml each.
This was a scenario based on assumed deployment locations, weather, grid water intensity and model energy. It was not a meter reading from ChatGPT. The paper’s value is showing that when and where a model runs can materially change its water footprint.
The order-of-magnitude gap with later provider figures may reflect more efficient hardware and cooling, different model energy, and narrower or simply different accounting. Because comparable underlying data is not public, assigning the gap to one cause would be speculation.
A simple calculation still needs hard inputs
At its simplest, direct water per query can be estimated as query energy multiplied by facility water-use effectiveness. Indirect water adds the energy multiplied by the water intensity of electricity generation. Training and manufacturing require separate allocations.
Each input moves. Query energy changes with model, input and output tokens, batching, hardware and tool calls. Facility water changes with weather and cooling mode. Grid water intensity changes with generation. An annual average smooths away the very hours when local water pressure may be highest.
Berkeley Lab modelled average US site water-use effectiveness at just over 0.36 litres per kWh in 2023. Multiplying that national fleet estimate by a product prompt estimate would mix unrelated systems, not reveal the provider’s actual water.
“Query” is not a stable unit
A short text completion might generate a few dozen tokens. A research request can process long documents, search the web and call a model several times. Image and video generation use different pipelines. Counting each as one query makes the denominator convenient but environmentally weak.
The better unit includes workload: model, input tokens, output tokens, modality and tool activity. Even then, shared serving infrastructure means allocation is approximate. Idle servers, networking and cooling support many users at once.
This is why how much water ChatGPT uses cannot be answered from message count alone. It is also why provider distributions would be more useful than one average: readers need to see ordinary, high-end and tool-heavy use separately.
Direct water is not the whole footprint
On-site consumption is locally important and often what permits regulate. It does not include water used by the power system. Conversely, a broad indirect model may assign water from distant generation that is less relevant to the data centre’s immediate catchment.
Chip factories use ultrapure water, and construction has supply-chain impacts. These stages are normally absent from per-query disclosures because no public method can confidently allocate a processor’s life across all its work. Training water has the same denominator problem.
Clear reporting presents several layers rather than claiming one is the truth: direct operational, indirect electricity, training allocation and embodied supply chain. Our guide to how AI uses water explains where each layer occurs.
How to quote a number responsibly
Name the service and statistic: “Google reported 0.26 ml for its median Gemini Apps text prompt in May 2025.” Preserve “median”, “text” and the date. For ChatGPT, say “Altman stated an average of about 0.32 ml, without publishing the method.”
When citing the academic work, say “the model estimated 500 ml across roughly 20–50 responses under its scenarios.” Do not shorten that to a bottle per message.
Treechat uses a fixed millilitres-per-watt-hour assumption to create consistent estimates from message energy. The methodology labels this as an estimate because actual cooling and power water are unavailable. Consistency can help compare Treechat messages; it cannot turn a generic factor into a facility measurement.
The practical takeaway is to reduce unnecessary computation and demand better disclosure. A transparent range is less memorable than a bottle, but much closer to what is actually known.