Does AI really use a bottle of water per message?
No: the original research estimated one 500 ml bottle for 20–50 responses, not one message. Newer figures are lower but count differently.

No. The research behind the claim did not say one AI message consumes a 500 ml bottle of water. It estimated that a bottle could correspond to roughly 20–50 medium-length language-model responses, depending on where and when the system ran. That is about 10–25 ml per response under the study’s assumptions.
Newer operator figures for ordinary text are lower still: about 0.26 ml for Google’s median Gemini Apps prompt and 0.32 ml for an average ChatGPT query stated by OpenAI’s chief executive. Those figures use different systems and methods, so they do not simply disprove the study. They do disprove presenting one bottle per message as a settled fact.
What the original paper actually said
The 2023 paper Making AI Less “Thirsty” set out to model the operational water footprint of large language models. It included on-site cooling and water associated with electricity generation, then tested how location and time changed the result.
Its memorable example estimated that GPT-3-class inference could consume a 500 ml bottle for roughly 20–50 responses. The range was essential: cooling weather and grid water intensity varied. It was not based on observing a bottle or dedicated pipe for one chat.
Some coverage dropped the denominator. “A bottle for a conversation” became “a bottle per prompt” or “per message”. That alteration multiplies the original estimate by 20–50 and removes the uncertainty the researchers were trying to highlight.
The distinction between a response and a conversation matters too. Conversations vary in length, and later turns may carry earlier context back into the model, changing the computation.
What newer providers report
Google reported a median Gemini Apps text prompt used 0.26 ml of water in a May 2025 point-in-time measurement. The study included accelerator, host and idle-machine energy plus data-centre overhead, then calculated water for its serving fleet.
Sam Altman stated an average ChatGPT query used 0.000085 US gallons, which converts to about 0.32 ml. OpenAI did not publish the calculation, request distribution or full water boundary, so the apparent precision should be treated cautiously.
Both are around a fraction of a millilitre, far below 10–25 ml. But a median and an average are different, and product infrastructure changed between the academic scenarios and 2025 disclosures.
Neither disclosure describes every image, file, search or reasoning mode. Applying a text-prompt figure to those workloads would repeat the same unit error in the opposite direction.
Why the estimates differ so much
The model and task affect electricity, which affects heat. The data-centre design affects direct water. Weather changes cooling demand. The power grid changes indirect water. Hardware, batching and utilisation change how much infrastructure is allocated to one response.
Boundaries may be the largest source of disagreement. Counting only operational site water is narrower than adding electricity generation. Allocating idle capacity, training or upstream manufacturing changes the answer again. Public disclosures do not offer enough matching detail to reconcile the figures precisely.
It would be equally wrong to choose the smallest number and declare AI water use solved. The spread shows that service-specific measurement is possible and that generic estimates need caveats. See AI water usage per query for a side-by-side explanation.
Operators could narrow the uncertainty by publishing ranges by model and workload, plus direct and indirect water for the facilities serving them.
A prompt does not literally consume water
Cooling systems serve racks and facilities continuously. Researchers divide a portion of consumption among workloads based on energy or activity. This allocation can support planning, but it is not a physical event triggered by the send button.
The same distinction applies to training. A training run may consume water at one time; analysts can distribute it across later messages, but the share depends on lifetime traffic. Chip-fabrication water is further upstream and seldom included.
These layers are laid out in how AI uses water. Keeping them separate prevents a direct cooling estimate from masquerading as a complete life cycle, while still recognising the real resources behind digital services.
It also keeps responsibility in view: facility operators control cooling and siting, while users mainly control the volume and type of work they request.
Bottle comparisons can still be useful
A familiar container makes an invisible volume understandable. The problem is not the metaphor; it is losing the source, denominator and boundary. “One bottle for 20–50 responses in a 2023 research scenario” is accurate. “Every message drinks a bottle” is not.
Ratios also become silly at very small volumes. At 0.26 ml, about 1,900 median Gemini prompts would mathematically equal 500 ml. At 0.32 ml, about 1,560 stated-average ChatGPT queries would. These are transparent divisions of published figures, not claims that those messages share one cooling system or environmental consequence.
Our everyday water comparisons preserve that distinction. A litre in a water-stressed catchment and a litre in a water-abundant one are equal volume but not equal impact.
What the real concern is
The environmental issue is aggregate infrastructure. Even a small per-query average can add to substantial demand when usage and computational intensity grow. Data centres also cluster, concentrating water and grid effects in particular communities.
Users can avoid unnecessary regeneration, limit irrelevant context and choose smaller models for routine work. Providers can do much more by publishing direct and indirect water, location, seasonal peaks and distributions by workload.
Treechat uses a fixed energy-to-water factor described on its methodology page. Its receipt is an estimate, not a live facility reading. That is the appropriate standard for all water claims: keep the method beside the number, and correct a memorable myth when it outruns its source.