Treechat
Menu
Essayer gratuitement
← All notes

Choosing an AI assistant by energy use

Choose a lower-energy AI assistant by comparing task success, model size, prompt and output length, system boundaries and provider disclosure.

By Treechat6 min read
A person compares several AI routes using capability, energy and transparency measures.

Choose the smallest AI system that reliably completes your task, then favour the provider with the clearest comparable energy evidence. Do not choose by model parameter count, a green badge or one company-wide renewable claim alone. The workload, hardware, utilisation, data-centre overhead and output quality all affect the result.

For most people, there is no audited energy label covering every commercial assistant on the same task. You may need to make a provisional choice: use ordinary search or software when it suffices, a small text model for routine drafting, and a larger reasoning or tool-using system only when added capability changes the outcome. Transparency is itself a selection criterion. A provider that states boundaries and uncertainty gives you more to assess than one offering an unexplained “eco” score.

Begin with the task, not the brand

Write down what success means before comparing assistants. Finding a known webpage, correcting spelling, summarising a report and debugging unfamiliar code require different capabilities. A fair comparison gives models the same input, asks for the same output length and checks whether the result is usable.

Energy without quality can reward systems that do less because they fail. If a small model needs repeated correction while a stronger model succeeds once, the lower-energy route is not obvious. Track retries, tool calls and generated tokens rather than counting only visible messages.

The reverse matters too. A frontier reasoning model may be excellent at difficult mathematics but unnecessary for rewriting a paragraph. Research by Luccioni, Jernite and Strubell found large differences between inference tasks and model types, including higher energy for general-purpose generative systems than task-specific models in their evaluated settings. Capability should be proportionate to the job.

Prefer comparable measurements

The strongest comparison measures models on the same hardware, software, task, input and output requirements. Public benchmarks such as the AI Energy Score use standardised tasks to make differences more visible. They are useful for the tested open or submitted models, not a complete ranking of consumer assistants whose backends can change.

Provider disclosures can add production reality. Google’s 2025 Gemini Apps inference study includes active accelerators, achieved utilisation, idle machines, CPU, memory and data-centre overhead. Its median prompt figure applies to Google’s defined product and period. It should not be placed beside a chip-only laboratory result as though the boundaries match.

When no comparable data exists, say so. Latency, subscription price and a laptop fan are not reliable energy proxies. A fast system may use more parallel hardware, while a slow one may simply be poorly utilised.

Read the system boundary

Ask whether an energy figure includes only accelerator power or the full server. Look for host CPU, memory, storage, networking, idle capacity and facility overhead. Data-centre PUE can add cooling and power-system overhead, but it cannot fill gaps inside the IT estimate.

Check whether the number covers inference only. Training and hardware manufacture are genuine impacts, yet allocating them to an individual answer depends on model lifetime and total use. A provider may report them separately rather than forcing them into an uncertain prompt figure.

Then examine carbon and water boundaries. Lower energy often helps both, but grid location and cooling design matter. Annual renewable matching does not reveal the electricity supplying a request hour by hour. A low PUE does not prove low water use. Avoid a single composite score unless the method explains its weighting and inputs.

A credible disclosure should let another reader reconstruct the comparison or at least identify every material assumption.

Look beyond the model name

An assistant may route easy questions to a small model and hard ones to a larger model. It may search the web, read files, call code tools or generate hidden reasoning before responding. One visible brand can therefore represent several workloads.

Controls matter. Prefer services that let you choose a model or reasoning level, limit output length and see when tools are used. Clear model labels are better than vague modes such as “smart” or “green” with no technical definition. A short receipt tied to the actual response is more informative than a lifetime average if its calculation is disclosed.

Context handling also changes the job. Continuing a conversation can be efficient when it avoids restating everything, but carrying a huge irrelevant history makes the model process more input. The lifecycle in what happens when you send an AI message explains why input and output tokens belong in any practical estimate.

Use a five-question shortlist

Before settling on an assistant, ask:

  1. Does the simplest non-AI tool already solve this task?
  2. What is the smallest model that produces an acceptable result without repeated retries?
  3. Does the provider publish task-specific energy with hardware, workload and system boundaries?
  4. Can I control model choice, context, output length and optional tools?
  5. Does the service distinguish energy, carbon, water and climate contributions rather than merging them into a green claim?

For a school, add privacy, accessibility and institutional policy. A student’s guide to AI environmental impact explains how to audit claims without turning resource use into guilt. For an organisation, test representative internal tasks and record success rate, response length and total request volume before procurement.

Your shortlist may have no measurable energy winner. In that case, choose on task fit, control and disclosure, then keep the decision open as better evidence arrives.

Treat rankings as snapshots

AI services change models, routing, hardware and software without always changing their product names. A result measured this month may not describe the same assistant next year. Record the test date, model version, feature settings and source link.

Efficiency per answer can also improve while total electricity rises because usage expands. The International Energy Agency’s Energy and AI scenarios make this system-level distinction clear. A provider should report absolute demand alongside intensity where possible.

Be cautious with claims that one assistant uses a fixed fraction of another’s energy. Unless both were tested on matched tasks and boundaries, the ratio is a modelled comparison, not a universal fact. This is the same problem behind unsupported ChatGPT-versus-Google-Search water comparisons.

Set a review date rather than treating the shortlist as permanent. New models can alter capability, routing and efficiency, while a familiar product name stays unchanged.

How Treechat fits the checklist

Treechat routes everyday questions towards a smaller model and uses a larger option when the task needs it. It estimates operational energy from model and token use, then adds a stated allowance for data-centre overhead and grid carbon. The methodology publishes the current assumptions and limits.

That receipt is an estimate, not direct measurement. It does not settle embodied hardware, training allocation or facility-specific water. Treechat’s tree-funding commitment is reported separately rather than subtracted from the estimated response footprint.

Those limits matter when comparing Treechat with a provider using a different system boundary or a directly measured fleet statistic.

The broader choosing rule remains provider-neutral: define the job, demand an acceptable answer, compare like with like and prefer the least computationally demanding route that works. If the evidence cannot identify a winner, do not invent one. Choose the service that gives you useful control and honest information, then reassess when comparable measurements become available.