Methodology

Our methodology

Every gram of CO₂ we report is backed by peer-reviewed research. Here's exactly how we calculate it — and where we're uncertain.

What we measure

We measure the operational emissions of the inference phase: the electricity a datacenter consumes to answer your requests.

Included

  • GPU/CPU electricity during inference
  • Datacenter overhead (cooling, distribution) via PUE
  • Regional grid carbon intensity

Not included

  • Model training (amortized across billions of requests)
  • Hardware manufacturing and transport
  • Your own device and network
  • Off-site cooling and water

These exclusions are consistent with Watershed, Climatiq, and most enterprise Scope 2 accounting practices. We state this scope explicitly so you can incorporate it correctly into CSRD reporting.

The calculation

CO₂ (gCO₂e) = tokens × (J / 1,000 tokens) × PUE × (gCO₂e / kWh) ÷ 3,600,000
  • tokens — input and output tokens are counted separately, since output generation is markedly more energy-intensive per token than reading input.
  • J / 1,000 tokens — the energy (in joules) a model consumes per 1,000 tokens processed, distinct for input vs. output.
  • PUE — Power Usage Effectiveness of the datacenter running the model (total facility power ÷ compute power).
  • gCO₂e / kWh — carbon intensity of the electricity grid powering that datacenter.
  • ÷ 3,600,000 — converts joules to kilowatt-hours (1 kWh = 3,600,000 J).

Input tokens are processed in parallel (prefill) — markedly less energy-intensive than output tokens, which are generated sequentially, one at a time. Not distinguishing the two understates the true carbon cost of long responses.

How we set the parameters

Energy per token

This is our main source of uncertainty. No provider publishes token-level consumption data. We rely on academic benchmarks run against real APIs under real deployment conditions — principally Jegham et al. (2025), “How Hungry is AI?”, University of Rhode Island, which benchmarked 30 commercial models for energy, water, and carbon at the prompt level. Values for a short query (100 input / 300 output tokens) range from about 0.1 Wh to 24 Wh depending on the model — a 240x spread.

ModelProviderPUEJ/1k inJ/1k outWh / short queryConfidence
GPT-4oopenai1.124003,5000.34medium
GPT-4o miniopenai1.121006,0000.56medium
GPT-4 Turboopenai1.1237018,7001.76medium
GPT-3.5 Turboopenai1.121507500.07low
GPT-4.1openai1.121,0839,4760.92medium
GPT-4.1 miniopenai1.12754,4860.42medium
GPT-4.1 nanoopenai1.12181,0970.10medium
o3-minireasoningopenai1.121009,0740.85medium
o4-mini (high)reasoningopenai1.1210031,2092.92medium
o1reasoningopenai1.1240047,5024.45medium
o3reasoningopenai1.1240075,1457.03medium
Claude Opus 5anthropic1.1490018,0001.74low
Claude Sonnet 5anthropic1.144409,0000.87low
Claude 3.5 Sonnetanthropic1.1486010,6001.03medium
Claude 3.7 Sonnetanthropic1.146958,5710.84medium
Claude 3.7 Sonnet (Extended Thinking)reasoninganthropic1.1469536,5053.49medium
Claude 3 Opusanthropic1.1480016,0001.55low
Claude Haiku 4.5anthropic1.1449013,6001.31low
Claude 3 Haikuanthropic1.143509,5000.91low
Gemini 1.5 Progoogle1.095102,0500.20medium
Gemini 1.5 Flashgoogle1.091305100.05low
Gemini 2.0 Flashgoogle1.091204800.05low
Mistral Largemistral1.204502,2000.23low
Mistral Smallmistral1.201105500.06low
DeepSeek-R1reasoningdeepseek1.27600224,82323.81medium
DeepSeek-V3deepseek1.2760033,0043.51medium
LLaMA 3.1 70Bmeta1.1535011,3721.10medium
LLaMA 3.1 8Bmeta1.151501,0250.10medium
GitHub Copilotopenai1.124003,5000.34low
Figma AIopenai1.124003,5000.34low
Notion AIopenai1.122752,1250.21low
Microsoft Copilotopenai1.1237018,7001.76low
Unknown / unlisted model—1.1345045000.44low

Emissio's emission factor database currently covers 29 AI models across 7 providers. The table above shows the subset with a fully physics-derived factor (energy per token, PUE, confidence level); the full model list — including cost estimates — is used across the product.

Power Usage Effectiveness (PUE)

PUE is the ratio of total datacenter energy to IT-equipment energy. A PUE of 1.12 means 12% of energy goes to cooling, power distribution, and other overhead beyond the compute itself.

InfrastructurePUESource
Microsoft Azure (OpenAI)1.12Microsoft Environmental Sustainability Report 2024
AWS (Anthropic)1.14AWS Sustainability Report 2024
Google Cloud (Gemini)1.09Google official fleet-wide average, datacenters.google/efficiency
Hyperscaler average (Mistral, LLaMA hosting)1.13Uptime Institute Global Data Center Survey 2024
China national average (DeepSeek)1.27China National Energy Administration 2024

Carbon intensity

Carbon intensity translates electricity into CO₂, based on the region's energy mix. We use location-based factors — the physical grid a datacenter draws from — not market-based factors, which fold in a provider's Power Purchase Agreements (PPAs) and reflect its climate commitments rather than the electricity it physically consumes. For rigorous CSRD reporting, the GHG Protocol recommends disclosing both — we report location-based.

RegiongCO₂e / kWhSource
US East (Virginia)269EPA eGRID SRVC subregion, 2023
US average386EPA eGRID national average, 2023
EU-27 average295ADEME 2024
France85ADEME 2024 (nuclear-heavy grid)
China600China National Energy Administration 2024
Global average494IEA 2024

The regions Emissio actually applies at calculation time are listed below — a conservative EU average is used when a datacenter's region is unknown.

Region codegCO₂e / kWh
EU295
FR85
DE385
GB233
US386
US-EAST415
US-WEST291
GLOBAL494
DEFAULT295

Confidence levels — what they actually mean

We show a confidence level for every model. It's not a rating of the model itself — it's our estimate of the uncertainty in the carbon number we calculate for it.

high
Official provider data with a published methodology, plus independent third-party verification of that exact model. Typical uncertainty: ±10-20%.
medium
A peer-reviewed academic benchmark measured against real APIs, or an official provider disclosure. Typical uncertainty: ±20-50%.
low
No infrastructure-level benchmark exists for this model — estimated by analogy to a measured model's architecture or size class. Typical uncertainty: ±50-70%.

No model is currently “high confidence.” No AI lab publishes per-token inference energy with a verifiable, real-time methodology. That gap is exactly why a page like this one needs to exist.

On Gemini specifically: Google disclosed in May 2025 that its median query consumes 0.24 Wh and emits 0.03 gCO₂. That second figure uses market-based factors that fold in Google's PPAs — it reflects Google's clean-energy commitments, not the physical carbon intensity of the grid it drew from. On a location-based basis, our calculation gives roughly 0.09 gCO₂ for the same median query.

Reasoning models — a different ballgame

“Reasoning” models (o3, o1, DeepSeek-R1, Claude Extended Thinking) generate a long internal chain of thought before answering. That multiplies output tokens — and the energy consumed — well beyond what a standard model uses for the same visible response.

Modelvs. standard siblingWh / short query
o3~17× GPT-4o7.03
o1~10× GPT-4o4.45
Claude 3.7 Sonnet (Extended Thinking)~4× Claude 3.7 Sonnet3.49
DeepSeek-R1~17× GPT-4o (+ China grid)23.81

Emissio detects reasoning models automatically and applies their distinct factors — tracking an o3 request as if it were GPT-4o would understate its footprint by roughly 17x. A short GPT-4o query costs about 0.34 Wh; the same prompt to o3 costs about 7.03 Wh.

Scope and exclusions

We don't calculate emissions from training. GPT-4-class models are estimated to have emitted on the order of hundreds of tonnes of CO₂e during training — but that cost is amortized across billions of requests and isn't triggered by your usage. We focus on what your AI consumption decisions actually affect: inference.

Similarly, your own hardware (laptop, internal servers) and network connection aren't included — they're shared with everything else you do and aren't specific to AI usage.

References

Frequently asked questions

Are your numbers certified?
No — no third party currently certifies AI inference carbon calculations. We apply the same methodology as the most recent peer-reviewed academic papers and publish our sources so they can be independently verified.
How do you handle models with no published data?
We estimate them by analogy to a measured model of similar architecture and size, and mark them “low confidence” with roughly ±60% uncertainty rather than presenting a false precision.
Why don't you use Google's official CO₂ figure for Gemini?
We do use it for energy (0.24 Wh per median query). For carbon, we prefer location-based factors over market-based factors that include PPAs, consistent with the GHG Protocol's Scope 2 guidance.
How often do you update these factors?
Whenever a major academic benchmark or official provider disclosure is published. Last updated: September 2026.

Last updated: September 2026