AI is becoming central to how companies write code, serve customers, create content, and run their businesses. Token counts and costs, even as they rise fast, are easy to track.
But one basic question is much harder to answer: How much energy — and ultimately how much carbon — does all of this AI actually use?
Today, we're introducing CLEER[Closed-model Latent Energy Estimation Range], an open, research-based approach designed to help answer that question.
AI is becoming central to how companies write code, serve customers, create content, and run their businesses. Token counts and costs, even as they rise fast, are easy to track.
But one basic question is much harder to answer: How much energy — and ultimately how much carbon — does all of this AI actually use?
Today, we're introducing CLEER[Closed-model Latent Energy Estimation Range], an open, research-based approach designed to help answer that question.
Open model energy usage can be directly measured with tools like AI Energy Score; for closed models (like GPT, Claude, Gemini), which are the most relevant to users today, this kind of measurement isn’t possible. CLEER combines the best of both worlds, testing open models under realistic conditions and applying these results to estimate closed model energy use.
Today, we’re releasing a Technical Report describing CLEER, a dataset representing realistic use cases, and a dashboard to visualize the data. We’re also publishing our approach to convert these energy factors to carbon emissions, based on cutting-edge research, standards, and methodologies.
This work is our best effort to characterize the environmental impacts of closed AI models based on thousands of experiments that we’ve run ourselves. We hope that this can pave the way for more disclosures from model developers, enabling customers to measure and manage these growing impacts and creating meaningful demand for more sustainable AI.
KEYTAKEAWAYS
1
1
For the same task, the choice of model can change energy use by >30X.
For the same task, the choice of model can change energy use by >30X.
2
2
Price is an imperfect guide to environmental impact: in about 1 in 5 head-to-head comparisons, the cheaper model had the higher estimated energy.
Price is an imperfect guide to environmental impact: in about 1 in 5 head-to-head comparisons, the cheaper model had the higher estimated energy.
3
3
For the same model, a typical agentic session used 27X more energy than a typical chat session, with a heavy agentic session reaching >100X more.
For the same model, a typical agentic session used 27X more energy than a typical chat session, with a heavy agentic session reaching >100X more.
4
4
Higher-capability model tiers (Opus/Pro/Sol) used nearly 4X as much energy, on average, as compared to faster, lightweight tiers (Haiku/Flash/Terra).
Higher-capability model tiers (Opus/Pro/Sol) used nearly 4X as much energy, on average, as compared to faster, lightweight tiers (Haiku/Flash/Terra).
5
5
Frontier models aren’t necessarily getting more efficient: release date shows no meaningful relationship with energy.
Frontier models aren’t necessarily getting more efficient: release date shows no meaningful relationship with energy.
How CLEER works
Step 01 / 04
Testing
CLEER directly measures widely used open models under realistic deployment conditions, varying hardware, batching, serving and parallelization configurations to capture how energy and performance change in practice.
20+
open models · 20B–2.8T parameters
2,400+
test runs
425M+
tokens measured
An estimate until there's something better
CLEER is designed to become more accurate as better data becomes available: since its estimates are calibrated against the limited provider disclosures that exist today, any new model-specific energy data can be incorporated directly as it emerges.
This means that if a model provider publishes trustworthy model-specific measurements, these can replace the corresponding CLEER estimates without requiring users to rebuild their accounting approach. CLEER fills the gap in the meantime, while also identifying exactly what kind of data providers would need to disclose to make environmental accounting more accurate.
Google’s Gemini models require a separate estimation approach, since Google serves these models on its proprietary Tensor Processing Units (TPUs), for which no accessible utility currently allows direct energy measurement. This means CLEER can’t build the same open model proxy set as calibrated for GPU-based models. Instead, Gemini estimates are anchored to Google’s published accelerator-energy disclosure and projected using observed Gemini performance, published TPU power characteristics, and assumptions about deployment configuration, with uncertainty propagated through Monte Carlo simulation (more detail here). [link forthcoming]
Putting the numbers to work
An approach built around hypothetical AI use can miss the messy questions that appear as soon as a company tries to apply it: Which usage data actually exists? How should model versions be handled? What happens when the provider or exact infrastructure is unknown? How should uncertainty flow through an inventory? And can the result become useful to engineering teams rather than remain a sustainability calculation performed once a year?
This is why we chose to develop CLEER alongside organizations confronting these questions in their own AI use and environmental accounting. We developed and tested this work with companies facing the problem in practice, including Etsy, Hg, and Sovos.
The pilots we ran with these companies helped turn theoretical questions into design requirements.
From carbon accounting to informing engineering decisions at Etsy
From carbon accounting to informing engineering decisions at Etsy
From carbon accounting to informing engineering decisions at Etsy
Etsy, the global marketplace for unique and creative goods, provides one of the clearest examples of where we want this work to go. First, we worked together with Etsy to develop and apply the approach to Etsy's AI-related carbon accounting: connecting token activity to accelerator energy, then carrying it through the additional infrastructure and emissions factors an inventory requires.
Then, instead of making this a calculation that’s done just once a year, we brought CLEER-derived environmental data into Etsy internal telemetry, so impact appears next to metrics like cost in the tools engineers already use. That approach makes it possible to inform decisions before the emissions happen.
Model routing is a great example. Directing a request to the smallest model capable of handling it well is an emerging sustainability practice, and its cost benefit is easy to demonstrate. Its environmental benefit has been much harder to quantify, because the energy difference between two closed models was not something anyone could put a number on. With CLEER per-model factors, that comparison becomes possible: routing decisions can be evaluated on carbon alongside cost and quality, and the environmental case for a change can be made in the same terms.
That builds on a pattern Etsy has been developing for years. Their earlier Cloud Jewels work made cloud resource use more visible and eventually contributed to the open-source Cloud Carbon Footprint project. The same idea now applies to AI: when engineers can see environmental impact alongside the other variables they already optimize, sustainability becomes part of everyday technical decision-making instead of a separate reporting exercise performed once a year.
"We’ve spent the past decade working to better understand and reduce the environmental impact of our digital world. AI is the latest chapter in that work. We wanted to support an approach that goes beyond Etsy, one the broader industry can align on, inspect, and improve. That’s why we helped test this methodology. It’s open, builds on past work such as Cloud Jewels, improves the accuracy of our disclosures, and gives our engineers something practical they can act on.”
— Chelsea Mozen, Head of Impact & Sustainability at Etsy
Etsy, the global marketplace for unique and creative goods, provides one of the clearest examples of where we want this work to go. First, we worked together with Etsy to develop and apply the approach to Etsy's AI-related carbon accounting: connecting token activity to accelerator energy, then carrying it through the additional infrastructure and emissions factors an inventory requires.
Then, instead of making this a calculation that’s done just once a year, we brought CLEER-derived environmental data into Etsy internal telemetry, so impact appears next to metrics like cost in the tools engineers already use. That approach makes it possible to inform decisions before the emissions happen.
Model routing is a great example. Directing a request to the smallest model capable of handling it well is an emerging sustainability practice, and its cost benefit is easy to demonstrate. Its environmental benefit has been much harder to quantify, because the energy difference between two closed models was not something anyone could put a number on. With CLEER per-model factors, that comparison becomes possible: routing decisions can be evaluated on carbon alongside cost and quality, and the environmental case for a change can be made in the same terms.
That builds on a pattern Etsy has been developing for years. Their earlier Cloud Jewels work made cloud resource use more visible and eventually contributed to the open-source Cloud Carbon Footprint project. The same idea now applies to AI: when engineers can see environmental impact alongside the other variables they already optimize, sustainability becomes part of everyday technical decision-making instead of a separate reporting exercise performed once a year.
"We’ve spent the past decade working to better understand and reduce the environmental impact of our digital world. AI is the latest chapter in that work. We wanted to support an approach that goes beyond Etsy, one the broader industry can align on, inspect, and improve. That’s why we helped test this methodology. It’s open, builds on past work such as Cloud Jewels, improves the accuracy of our disclosures, and gives our engineers something practical they can act on.”
— Chelsea Mozen, Head of Impact & Sustainability at Etsy
Making AI emissions visible inside one company is useful. Scaling that across the systems companies already use to manage their carbon footprints is where this approach can have much broader impact.
That is why we’re excited to announce a new design partnership with Watershed, the enterprise sustainability platform used by more than 800 global companies. In July, Watershed published its open framework for AI emissions measurement, a tiered approach designed to help companies account for AI emissions with the data available to them today and improve those estimates as better data becomes available.
SAIG was pleased to participate in the review process and provide feedback on the framework, which also builds on prior research from members of our team. The work helped move corporate AI emissions accounting forward while underscoring one of the field’s biggest remaining data gaps: model-specific energy intensity. Watershed’s framework specifically calls for model-specific values where possible and identifies per-token energy intensity as a key input to its activity-based approach.
Our design partnership is intended to help close that gap. SAIG and Watershed will explore how CLEER’s model-specific factors can inform Watershed’s activity-based AI emissions accounting and what would be required to make those factors usable within carbon accounting systems.
At SAIG, our goal is to make AI emissions accounting as easy to use as possible: in the flow of work, using the data companies already have, and inside the systems where they already manage their emissions. Measurement matters most when it leads to action. By connecting more granular model-level data with established carbon accounting infrastructure, we hope to make AI sustainability not just measurable, but actionable.
“Our AI emissions measurement framework identified a critical gap: reliable, model-specific data for closed AI models. We’re excited to work with SAIG to explore how that gap can be closed and turn better measurement into action on AI sustainability.”
Making AI emissions visible inside one company is useful. Scaling that across the systems companies already use to manage their carbon footprints is where this approach can have much broader impact.
That is why we’re excited to announce a new design partnership with Watershed, the enterprise sustainability platform used by more than 800 global companies. In July, Watershed published its open framework for AI emissions measurement, a tiered approach designed to help companies account for AI emissions with the data available to them today and improve those estimates as better data becomes available.
SAIG was pleased to participate in the review process and provide feedback on the framework, which also builds on prior research from members of our team. The work helped move corporate AI emissions accounting forward while underscoring one of the field’s biggest remaining data gaps: model-specific energy intensity. Watershed’s framework specifically calls for model-specific values where possible and identifies per-token energy intensity as a key input to its activity-based approach.
Our design partnership is intended to help close that gap. SAIG and Watershed will explore how CLEER’s model-specific factors can inform Watershed’s activity-based AI emissions accounting and what would be required to make those factors usable within carbon accounting systems.
At SAIG, our goal is to make AI emissions accounting as easy to use as possible: in the flow of work, using the data companies already have, and inside the systems where they already manage their emissions. Measurement matters most when it leads to action. By connecting more granular model-level data with established carbon accounting infrastructure, we hope to make AI sustainability not just measurable, but actionable.
“Our AI emissions measurement framework identified a critical gap: reliable, model-specific data for closed AI models. We’re excited to work with SAIG to explore how that gap can be closed and turn better measurement into action on AI sustainability.”
CLEER is being released through a public dashboard that translates the approach into representative chat and agentic use cases, using standardized workloads derived from public production datasets. The dashboard brings energy, cost and performance together so models can be evaluated more holistically, with tools to build custom visualizations, compare models against one another, and examine how a single model performs across different types of tasks. This dataset is available via GitHub today, and API access is planned in the near future.
For organizations and software providers that need model-specific factors for carbon inventories, engineering telemetry, or other operational tools, CLEER Feed provides model-level data for integration. Reach out to learn more.
Feedback and contributions
CLEER and the accompanying emissions calculation are intended to improve through use, feedback, and new evidence. We welcome technical feedback, proposed corrections, and additional data. The best way to contribute is to open an issue (here for CLEER Technical Report, here for Closed AI Emissions Calculation) on the relevant GitHub repository, so discussion and revisions remain transparent and traceable. We also intend to submit CLEER for formal peer review.
What comes next
CLEER is a starting point for estimating AI’s environmental impacts, not a substitute for transparency.
As serving architectures evolve, our benchmarking protocols will evolve with them. As providers disclose more information, we will incorporate it into our approach. As enterprises apply the approach, we expect to learn where accounting guidance needs to become more precise and keep improving it.
Most importantly, we want these numbers to become useful and actionable.
If companies better understand the environmental impacts of the models they deploy, sustainability can become another dimension of their decision-making, alongside performance, speed, and price.
That’s the larger objective behind this work: to make AI's environmental differences visible, put that information into the hands of the organizations buying AI, and create stronger incentives for the industry to build more sustainable systems.
Acknowledgements
This work has benefited from companies willing to test the methodology in real-world settings and partners that contributed the compute, data and technical support required to build it.
Acknowledgements
This work has benefited from companies willing to test the methodology in real-world settings and partners that contributed the compute, data and technical support required to build it.
Acknowledgements
This work has benefited from companies willing to test the methodology in real-world settings and partners that contributed the compute, data and technical support required to build it.
Pilot Partners
We thank Etsy, Hg, and Sovos for participating as pilot customers and contributing practical feedback that helped us test and refine the methodology in real-world enterprise contexts.
Compute and Technical Partners
Pilot Partners
We thank Etsy, Hg, and Sovos for participating as pilot customers and contributing practical feedback that helped us test and refine the methodology in real-world enterprise contexts.
Compute and Technical Partners
Pilot Partners
We thank Etsy, Hg, and Sovos for participating as pilot customers and contributing practical feedback that helped us test and refine the methodology in real-world enterprise contexts.
Compute and Technical Partners
This work was supported by Radium, whose technical team provided engineering support and access to the resources required to conduct the experimental evaluation.
This work was supported by Lambda, which provided compute resources used in the experimental benchmarking underpinning the methodology.
Neuralwatt, an AI inference platform focused on measuring and optimizing energy use, provided open-weight model data and feedback that helped us characterize request patterns and calibrate the methodology alongside SAIG's own benchmarking.