The Most Sustainable AI Model: Comparing Carbon Footprint, Energy and Water Use
Determining the 'cleanest' or 'most sustainable' AI model requires separating two distinct variables: hardware energy efficiency, measured in watt-hours per token, and the carbon intensity of the grid powering that compute, measured in CO₂e per kWh. When applying Lifecycle Assessment (LCA) methodology, the operational footprint is primarily determined by infrastructure utilisation rates and the physical location of data centres — not by model brand. So how do Google Gemini, OpenAI's GPT, Anthropic's Claude, and Mistral compare?
Based on current environmental benchmarking and corporate disclosures, Google Gemini currently reports the lowest verified Global Warming Potential (GWP) per standard query. EU-based Mistral presents a highly competitive alternative depending on deployment configuration. OpenAI's reasoning models carry a materially higher per-query footprint. Anthropic has not published comparable metrics — and a recent infrastructure deal with xAI for computer raises new questions about Claude's emissions trajectory. This page covers what sustainability professionals and environmentally-conscious individuals need to know when accounting for AI's climate footprint and GHG inventory.

AI data centres use the resources available to them — including energy from the local grid. An AI data centre could be powered entirely be renewable energy, store, and other sustainable sources of energy, or by coal, natural gas, and other polluting fossil fuels
What 'Clean' or 'Sustainable' Means for an AI Model
Sustainable in the context of AI use has two components that move independently. A model can be computationally efficient — requiring few watt-hours per token — but running on a coal-heavy grid will produce high carbon output. Conversely, a less efficient model hosted in a low-carbon, renewable energy-powered data centre may produce lower net emissions per query. Comparing providers requires tracking both.
Lifecycle Assessment methodology for AI systems distinguishes between:
- Operational carbon: electricity consumed during inference, multiplied by the grid's carbon intensity at the time and location of compute
- Embodied carbon (Scope 3): emissions from manufacturing the hardware itself — GPUs, servers, cooling systems — amortised across the operational lifetime of the infrastructure
- Water consumption: indirect environmental impact from evaporative cooling systems. Not a direct GHG metric but increasingly material for climate impacts and sustainability
Most published AI sustainability figures report only operational carbon. For full-scope reporting under the GHG Protocol or CSRD/ESRS E1, all three categories are relevant. See our Scope 3 categories guide for how AI compute is typically classified.
How Major AI Providers Compare on Operational Emissions
The figures below reflect operational carbon only — energy consumed during inference multiplied by the relevant grid carbon intensity. Embodied carbon and water consumption are addressed separately below.
Google Gemini
Google operates the most optimised large-scale inference infrastructure currently benchmarked. Published disclosures indicate a median text prompt consumes approximately 0.24 Wh, yielding a Global Warming Potential of roughly 0.03 grams CO₂e. This efficiency is driven by Google's Carbon-Free Energy (CFE) programme, which averages 64% globally and exceeds 90% in select availability zones.
Google's CFE methodology matches clean energy procurement to the specific hour and location of compute — a more rigorous approach than annual renewable energy certificates (RECs), which are the standard for most cloud providers. This temporal and geographic matching is why Gemini's per-query GWP is significantly lower than its raw energy consumption figure alone would suggest.
OpenAI (ChatGPT / GPT-4o)
An average standard GPT-4o query consumes approximately 0.34 Wh, resulting in roughly 0.165 grams CO₂e based on a blended US grid mix. This is higher than Gemini on both the energy and carbon dimensions, reflecting different infrastructure efficiency and a lower CFE procurement rate.
The more significant concern for sustainability teams is OpenAI's reasoning model line — o1, o3, and successors. These models execute internal reasoning chains before producing an output, generating thousands of hidden tokens per user request. A single o3 query may consume five to ten times the energy of an equivalent GPT-4o request, depending on the complexity of the reasoning chain triggered. For organisations logging AI compute in their GHG inventory, the specific model tier used — not just the provider — is a material input to the emissions calculation.
Mistral
As an open-weights model provider, Mistral's environmental footprint depends entirely on the inference engine and hosting environment. Earlier benchmarks placed Mistral's hosted web footprint higher — approximately 1.14 grams CO₂e per query — but this reflected sub-optimal deployment configurations. 2025 benchmarks using vLLM inference engines on H100 GPUs show materially lower energy draws: approximately 0.022 Wh of GPU energy for 400 output tokens on Mistral Small 3.2.
The dominant variable for Mistral is grid location. Mistral's primary data centre infrastructure is in France, where the electricity grid is heavily decarbonised by nuclear generation — approximately 57 grams CO₂/kWh, compared to a US average of 386 grams CO₂/kWh. This five-to-tenfold difference in grid carbon intensity means that a Mistral workload routed through French infrastructure may carry lower operational carbon than a more computationally efficient model running in a coal-heavy US availability zone. Grid emissions factors are the dominant variable for Mistral deployments — not model-level efficiency.
Anthropic (Claude)
Anthropic has not publicly disclosed per-token energy consumption, watt-hours per query, or Global Warming Potential metrics for any Claude model. Direct comparison on a like-for-like basis with Google or OpenAI is not currently possible.
Claude runs on AWS and Google Cloud infrastructure, with availability zone selection determined by Anthropic's deployment architecture. Carbon intensity fluctuates significantly across AWS and GCP zones — ranging from under 50g CO₂/kWh in some regions to over 500g CO₂/kWh in others. Without disclosure of which zones serve inference traffic, and at what times, the operational carbon of a Claude query cannot be independently verified.
Anthropic's Colossus Deal: A Material Climate Risk
In 2025, Anthropic signed an agreement to use computing capacity at Colossus 1, a large-scale data centre in Memphis, Tennessee operated by xAI. The arrangement is intended to increase capacity for paid Claude users.
Colossus launched in 2024 with over 35 methane-spewing gas turbines as primary power infrastructure. A significant number of these turbines were classified as 'temporary' equipment — a classification that allowed xAI to avoid the Clean Air Act permitting requirements that would otherwise apply to a facility of this scale. Environmental regulators and local health advocates subsequently contested this classification.
The facility has been identified as one of the largest industrial emitters of nitrogen oxides in the Memphis metropolitan area. Residents near the site have documented increased days per year when outdoor air quality falls below safe thresholds. These permitting disputes and local health impacts were active and unresolved at the time Anthropic announced its capacity deal.
For organisations with supplier or value chain emissions obligations — particularly under CSRD/ESRS E1, CSDDD, or SBTi supply chain requirements — the upstream infrastructure decisions of AI providers constitute a Scope 3 exposure. Where assurance is required on ESG data quality, material changes to a supplier's environmental profile — such as a significant new infrastructure dependency — should be assessed, documented, and where material, disclosed.
Beyond Per-Query Carbon: Full Lifecycle Footprint Factors
Operational carbon represents only a portion of the total environmental footprint of AI compute. A rigorous sustainability disclosure requires accounting for three additional dimensions.
Embodied Carbon (Scope 3, Category 1 or 2)
A significant share of an AI model's lifecycle emissions comes from manufacturing the hardware on which it runs — GPUs, servers, cooling systems, and data centre construction. These emissions are incurred before the hardware is ever powered on. For environmental disclosures, the amortised embodied carbon of the server infrastructure — derived from verified Environmental Product Declarations (EPDs) where available — must be allocated across every token generated over the hardware's operational lifetime.
A key implication: a computationally efficient model running on underutilised infrastructure carries a higher embodied carbon cost per token than a less efficient model running at high utilisation. The numerator (total hardware embodied carbon) is fixed by manufacturing; the denominator (tokens generated) scales with utilisation. Providers with high GPU utilisation rates spread that fixed cost across a larger output volume.
Water Consumption
High-density compute relies on evaporative cooling, which consumes fresh water — a resource under increasing physical and regulatory scrutiny in many regions. Google and OpenAI average between 0.26 mL and 0.32 mL of water consumed per query at their optimised data centre facilities. Less optimised third-party hosting environments have been benchmarked at over 40 mL per query — more than 100 times higher.
Water consumption does not appear in GHG inventories, but it is increasingly material to CSRD/ESRS E3 (Water and Marine Resources) disclosures and to CDP Water questionnaire responses. Organisations reporting on AI-related water use require disclosure from their AI providers — which most currently do not provide at a per-customer or per-workload level.
Model Rightsizing
Deploying a frontier model with hundreds of billions of parameters to perform simple data extraction or classification wastes compute and raises emissions unnecessarily. Matching model size to task complexity is consistently the most effective lever for reducing AI compute emissions — more effective than switching providers.
- Simple extraction and classification: models with 7–13B parameters are typically sufficient
- Structured data generation and summarisation: 30–70B range provides quality without frontier-scale energy cost
- Complex reasoning, synthesis, and novel problem-solving: frontier models are justified; avoid reasoning-chain variants unless the task genuinely requires multi-step inference
- Batch processing: schedule asynchronous workloads to run during low-carbon grid hours where provider scheduling supports it
What This Means for Your GHG Inventory
For most organisations, AI compute sits in Scope 2 (if running on owned or leased infrastructure) or Scope 3 Category 1 (purchased goods and services) if consumed via API. The correct classification depends on the contractual relationship and degree of operational control. A well-documented GHG Inventory Management Plan should specify the methodology used to estimate AI compute emissions: the data source (provider disclosures, energy bills, or spend-based proxy), the emissions factor applied, and the update cadence.
Where provider disclosures are unavailable or insufficiently granular — as is currently the case for Anthropic — a spend-based estimation approach using economy-wide spend factors from the EEIO model is the fallback method accepted by the GHG Protocol. The absence of provider data is not an acceptable reason to exclude AI compute from a reported inventory. For companies subject to GHG assurance, auditors will ask how material AI compute expenditure has been handled — and 'not included' is a finding.

Need Support with Sustainable AI?
Brightest helps teams implement responsible AI for sustainability, risk management, and performance
