AI has an environmental cost.
See yours.

Every conversation, image and video draws power from the grid. We think you deserve to know that, so EchoChat shows the energy behind every reply, before you send it.

What one reply costsShown live in-app
0.24Wh

Quick chata short everyday question, about two minutes of an LED bulb.

Quick chatLong chatDeep thinkingMedia
Every model
carries a leaf efficiency rating
Per reply
energy shown before you hit send

Why AI energy is hard to measure

Almost no major AI provider publicly discloses per-request energy data. Google is the only large provider to have published figures. Gemini uses roughly 0.24 Wh per median text query, and OpenAI's CEO has cited about 0.34 Wh per ChatGPT query. Anthropic, Meta, Mistral, DeepSeek and xAI have disclosed nothing.

10/13

major AI companies disclose zero data on energy, carbon or water use, per the Stanford Foundation Model Transparency Index.

The data isn't missing because it doesn't exist. Companies simply aren't required to share it, and many treat it as commercially sensitive.

Who publishes their energy use?
Per-request figures, publicly disclosed.
  • Google~0.24 WhPublished
  • OpenAI~0.34 WhOffhand remark
  • AnthropicNothing
  • MetaNothing
  • MistralNothing
  • xAI & othersNothing
3/13

major providers have said anything at all. EchoChat estimates the rest, so your team is never in the dark.

What drives energy use

Not all AI interactions are equal. Four factors determine the footprint.

8B
70B
235B
Model size

A model with 8 billion parameters requires substantially less energy than one with 200 billion. Smaller models like Haiku and Flash are specifically designed for efficiency.

Dense
MoE
Architecture

Mixture of Experts (MoE) models activate only 10-20% of parameters per request. A 235B MoE model may use less energy than a 70B dense model that activates everything.

standard
10–50×thinking mode
Reasoning modes

Extended thinking generates hidden intermediate computation. This multiplies energy by 10-50x compared to a standard reply. The most consequential variable to understand.

Text
Image
Video
Media generation

Generating a single high-resolution image or video clip is orders of magnitude more energy-intensive than a text conversation. Video generation sits at the high end.

T(h)ree commitments

Not greenwashing pledges. Concrete product decisions you can see and verify.

Show the footprint

Every model in EchoChat carries an efficiency rating, from one to five leaves, based on active parameter count, architecture type, and whether the model uses extended reasoning. You can see it before you send.

Default to efficient

The default text model in EchoChat is Mistral, a European model with published efficiency benchmarks. It is a capable model, and the responsible starting point.

Local models: no external compute

With EchoChat's local model tier, inference runs on hardware inside your workspace. No external API call, no external energy consumption. For high-volume teams, local models are both the most private and the most efficient option.

Running a large frontier model requires significant compute. A single inference request from a large model can consume more energy than hundreds of requests from a small efficient one, for tasks where the smaller model performs equally well.

The industry default is to reach for the most capable model available. Almost no AI product tries to address this: no indicator, no default toward efficiency, no option to stay local. We think that is a mistake, ethically and practically.

Teams that develop good intuitions about which model to reach for, and why, get better outcomes at lower cost with less waste. The efficiency rating is a nudge toward that intuition, not a restriction.

Thinking mode multiplier
The biggest energy variable in the catalog.
Standard reply
10–50×Thinking mode
energy per interaction, same prompt

What "thinking mode" costs

When a model thinks through a problem step by step before responding, it generates large amounts of intermediate computation not visible in the final reply. This can multiply energy consumption by 10 to 50 times compared to a standard reply.

Models like o3, DeepSeek R1, and the thinking variants of Gemini and Qwen fall into this category. They are capable tools, but they come with a proportionally larger energy cost. EchoChat shows a clear indicator whenever thinking mode is active, so users can make an informed choice.

Figures last reviewed April 2026. Sources include Google DeepMind, Epoch AI, the Green Software Foundation, Stanford CRFM, and TokenPowerBench.

Choose AI that counts its own footprint.

Default to efficient. Make the cost visible. Choose local when it counts.