An Open Tool for Estimating the Environmental Footprint of LLMs

Or click an example to test it

Visualisation

Comparative environmental impact of models

This public benchmark currently quantifies 83 market models. 56 use a strict parameter basis (directly sourced, directly derivable, or sourced from an explicit third-party estimate), while 27 additional recent models use a documented partial-data donor prior derived from the strict catalog.

Retained inference scenario1 hour of active use

The chart below shows the estimated central values for the 83 quantified catalog models under a standardized inference scenario corresponding to 1 hour of active use: 34.6 interactions/hour, 1000 input tokens, 550 output tokens, and one LLM request per use. The hourly pace is derived from an average reading speed of 238.0 words/min (Brysbaert, 2019) and a project convention of 1 token ≈ 0.75 word.

Benchmarks integrated into the chart now include direct household-use examples from Purdue Extension (fluorescent lamp ≈ 9.3 Wh over 1 h; laptop ≈ 32 Wh over 1 h), plus explicit project-scale equivalents for everyday uses better aligned with the current LLM range: laptop use for 3 h96 Wh, laptop use for 5 h160 Wh, electric kettle for 7 min210 Wh, and electric space heater for 10 min250 Wh. For carbon, the chart uses an average gasoline car benchmark derived from the ICCT (2025) factor retained by the project (235 gCO2e/km), shown here both for 430 m101.1 gCO2e and for 1 km235 gCO2e. Exact retained benchmark values are listed in the annex.

Trade-off

Inference vs. training impact map

This scatter plot compares each quantified model on two axes at once: standardized inference impact over one hour on the horizontal axis and retained training impact on the vertical axis. Point size follows the retained active parameter basis, while colors distinguish providers.

Positioning

Inference bubble chart

This explicit bubble chart positions each model by effective active parameters and by its retained inference impact. Bubble size reflects the retained context window, while colors distinguish providers.

Uncertainty

Inference uncertainty span by model

This chart makes the project’s retained inference range explicit for each model by showing the low, central, and high values under the standardized one-hour scenario.

Positioning

Inference model landscape

This landscape view clusters the quantified models from the characteristics retained by the project for inference screening: active and effective parameter basis, context window, serving mode, modality support, architecture notes, and central energy and carbon outputs. Nearby points indicate models with similar retained screening profiles, not a simple one-metric ranking.

Proxy

Inference screening factor heatmap

This heatmap exposes the central screening factors retained for each quantified market model. It shows the four multiplicative factors used by the project’s prompt proxy and the resulting ratio between effective and raw active parameters.

Positioning

Inference carbon vs. parameter count

This complementary view places models by retained active parameter count on the horizontal axis and by central inference carbon over one hour on the vertical axis, using logarithmic scaling on both axes.

Sensitivity

Country-mix sensitivity

This view compares central inference energy and carbon over one hour, while coloring each model by the retained electricity-mix country used for carbon recalculation. It helps separate model-size effects from country-mix effects.

Timeline

Inference carbon by model release date

This timeline follows the evolution of the project’s central inference CO2e estimate over time for the OpenAI, Claude, Grok, and Mistral families, using the release month of each model as the horizontal axis.

Perspective

Inference CO2 doubling view

This discussion-oriented chart summarizes the central inference screening values for flagship GPT, Claude, and Grok models as a simple doubling-time reading. It should be read as an interpretation of the retained observatory values, not as a provider-side measurement law.

Under the current central screening profile, the flagship inference series suggests a slower increase than training, because standardized usage, active compute, and per-request serving assumptions damp part of the growth that appears in total model scale.

Visualisation

Comparative training impacts of models

This public training benchmark currently quantifies 83 market models: 56 with a strict retained parameter basis and 27 with a documented partial-data donor prior.

The chart below shows the central values retained for the quantified models across two training indicator families: training energy and training CO2e contextualized from retained training energy and the model country proxy. The current screening method combines retained parameter count, a training-token prior, a training-regime prior, architecture features, and a hardware-class proxy. Everyday benchmarks are inserted directly into the list to situate those scales, not to imply direct observed equivalence.

Benchmarks integrated into the chart: household electricity for 74,277 households over one year of domestic use, i.e. ≈ 0.19 TWh based on an average consumption of 2,500 kWh per household (RTE, 2021 estimate), and full-flight aviation derived from Klöwer et al. (2025) from 577.97 MtCO2 and 27.45 million commercial flights observed in 2023, i.e. ≈ 71,491.6 tCO2e for 3,396 full flights. These comparison points are aligned with the current maximum central screening order of magnitude in the training chart, not with a direct provider-side measurement.

Positioning

Training model landscape

This landscape view clusters the quantified models from the characteristics retained by the project for training screening: retained parameter basis, training-token prior, training regime, hardware-class proxy, modality support, architecture notes, and central training energy and carbon outputs. Nearby points indicate similar retained screening profiles rather than a direct ranking on one axis.

Positioning

Training screening factor heatmap

This heatmap exposes the central screening factors retained for each quantified market model in the training proxy. It shows the regime, architecture, and hardware factors together with the retained training-token ratio per parameter.

Uncertainty

Training uncertainty span by model

This view shows the low, central, and high training CO2e values contextualized from retained training energy for each quantified market model. It makes explicit how widely the training proxy can vary once the parameter and token exponents, donor priors, and contextual factors are widened.

Positioning

Training carbon vs. parameter count

This complementary view places models by retained parameter count on the horizontal axis and by contextualized training CO2e on the vertical axis, using logarithmic scaling on both axes.

Timeline

Training CO2e by model release date

This timeline follows the evolution of the project’s retained training CO2e estimate over time for the OpenAI, Claude, Grok, and Mistral families, using the release month of each model as the horizontal axis.

Perspective

Training CO2e doubling view

This chart compresses the central training screening values of flagship GPT, Claude, and Grok models into a simple doubling-time interpretation. It is meant as a discussion support to make structural acceleration legible, not as a claim of direct industrial telemetry.

The apparent acceleration is stronger for training because the current screening method compounds retained parameter count, token priors, architecture effects, and hardware assumptions. The resulting doubling pace is therefore a transparent scenario reading, not a universal empirical constant.

Models

83 market models quantified by the project

The table below compares the market models retained in the public quantitative benchmark under the same inference scenario. For each model, the application shows the central values produced by the project’s multi-factor prompt proxy, both per hour of standardized use and per request. The current catalog combines 56 strict source-linked rows and 27 partial-data rows derived from donor models with comparable sourced metadata.

100B* US
Screening proxy
6.0 Wh 2.3 gCO2e 0.17 Wh 0.0669 gCO2e
100B* US
Screening proxy
6.0 Wh 2.3 gCO2e 0.17 Wh 0.0669 gCO2e
100B* US
Screening proxy
6.0 Wh 2.3 gCO2e 0.17 Wh 0.0669 gCO2e
100B* US
Screening proxy
6.0 Wh 2.3 gCO2e 0.17 Wh 0.0669 gCO2e
22B* US
Screening proxy
1.4 Wh 0.55 gCO2e 0.0412 Wh 0.0159 gCO2e
175B* US
Screening proxy
10.2 Wh 3.9 gCO2e 0.30 Wh 0.11 gCO2e
175B* US
Screening proxy
10.2 Wh 3.9 gCO2e 0.30 Wh 0.11 gCO2e
5000B* US
Screening proxy
264.7 Wh 101.9 gCO2e 7.7 Wh 2.9 gCO2e
2000B* US
Screening proxy
103.5 Wh 39.9 gCO2e 3.0 Wh 1.2 gCO2e
5000B* US
Screening proxy
264.7 Wh 101.9 gCO2e 7.7 Wh 2.9 gCO2e
4000B* US
Screening proxy
214.1 Wh 82.4 gCO2e 6.2 Wh 2.4 gCO2e
175B* US
Screening proxy
10.2 Wh 3.9 gCO2e 0.30 Wh 0.11 gCO2e
400B* US
Screening proxy
24.0 Wh 9.2 gCO2e 0.69 Wh 0.27 gCO2e
22B FR
Provider-country proxy
1.2 Wh 0.0484 gCO2e 0.0347 Wh 0.0014 gCO2e
37B active / 671B total CN
Provider-country proxy
2.9 Wh 1.6 gCO2e 0.0836 Wh 0.0451 gCO2e
37B active / 671B total CN
Provider-country proxy
2.8 Wh 1.5 gCO2e 0.0798 Wh 0.0431 gCO2e
13B active / 284B total CN
Provider-country proxy
1.1 Wh 0.61 gCO2e 0.0327 Wh 0.0177 gCO2e
49B active / 1600B total CN
Provider-country proxy
4.3 Wh 2.3 gCO2e 0.12 Wh 0.0675 gCO2e
123B FR
Provider-country proxy
6.8 Wh 0.27 gCO2e 0.20 Wh 0.0078 gCO2e
133.5B* US
Screening proxy
8.5 Wh 3.3 gCO2e 0.25 Wh 0.0944 gCO2e
133.5B* US
Screening proxy
8.5 Wh 3.3 gCO2e 0.25 Wh 0.0944 gCO2e
30B* US
Screening proxy
2.1 Wh 0.79 gCO2e 0.0594 Wh 0.0229 gCO2e
80B* US
Screening proxy
5.2 Wh 2.0 gCO2e 0.15 Wh 0.0581 gCO2e
30B* US
Screening proxy
2.1 Wh 0.79 gCO2e 0.0594 Wh 0.0229 gCO2e
200B* US
Screening proxy
12.5 Wh 4.8 gCO2e 0.36 Wh 0.14 gCO2e
30B* US
Screening proxy
2.1 Wh 0.79 gCO2e 0.0594 Wh 0.0229 gCO2e
3000B* US
Screening proxy
163.2 Wh 62.8 gCO2e 4.7 Wh 1.8 gCO2e
30B* US
Screening proxy
1.9 Wh 0.72 gCO2e 0.0543 Wh 0.0209 gCO2e
22B* US
Screening proxy
1.3 Wh 0.49 gCO2e 0.0369 Wh 0.0142 gCO2e
22B* US
Screening proxy
1.5 Wh 0.59 gCO2e 0.0442 Wh 0.0170 gCO2e
3000B* US
Screening proxy
163.2 Wh 62.8 gCO2e 4.7 Wh 1.8 gCO2e
130B US
Documented region, country retained as reference
6.9 Wh 2.7 gCO2e 0.20 Wh 0.0768 gCO2e
280B GB
Documented region, country retained as reference
14.3 Wh 2.6 gCO2e 0.41 Wh 0.0744 gCO2e
175B* US
Screening proxy
9.2 Wh 3.5 gCO2e 0.26 Wh 0.10 gCO2e
440B* active / 1760B* total US
Screening proxy
26.0 Wh 10.0 gCO2e 0.75 Wh 0.29 gCO2e
200B* US
Screening proxy
11.4 Wh 4.4 gCO2e 0.33 Wh 0.13 gCO2e
200B* US
Screening proxy
11.4 Wh 4.4 gCO2e 0.33 Wh 0.13 gCO2e
95B* US
Screening proxy
5.9 Wh 2.3 gCO2e 0.17 Wh 0.0657 gCO2e
95B* US
Screening proxy
5.9 Wh 2.3 gCO2e 0.17 Wh 0.0657 gCO2e
3000B* US
Screening proxy
156.8 Wh 60.4 gCO2e 4.5 Wh 1.7 gCO2e
440B* US
Screening proxy
25.3 Wh 9.7 gCO2e 0.73 Wh 0.28 gCO2e
3000B* US
Screening proxy
156.8 Wh 60.4 gCO2e 4.5 Wh 1.7 gCO2e
95B* US
Screening proxy
5.9 Wh 2.3 gCO2e 0.17 Wh 0.0657 gCO2e
95B* US
Screening proxy
5.9 Wh 2.3 gCO2e 0.17 Wh 0.0657 gCO2e
3000B* US
Screening proxy
156.8 Wh 60.4 gCO2e 4.5 Wh 1.7 gCO2e
3000B* US
Screening proxy
156.8 Wh 60.4 gCO2e 4.5 Wh 1.7 gCO2e
3000B* US
Screening proxy
163.3 Wh 62.9 gCO2e 4.7 Wh 1.8 gCO2e
5.1B active / 117B total US
Comparative reference country
0.40 Wh 0.16 gCO2e 0.0116 Wh 0.0045 gCO2e
3.6B active / 21B total US
Comparative reference country
0.26 Wh 0.10 gCO2e 0.00740 Wh 0.0029 gCO2e
78.5B* active / 314B* total US
Documented region, country retained as reference
5.5 Wh 2.1 gCO2e 0.16 Wh 0.0612 gCO2e
78.5B* US
Documented region, country retained as reference
4.8 Wh 1.8 gCO2e 0.14 Wh 0.0531 gCO2e
115B* active / 270B* total US
Documented region, country retained as reference
8.3 Wh 3.2 gCO2e 0.24 Wh 0.0920 gCO2e
600B* US
Documented region, country retained as reference
38.0 Wh 14.6 gCO2e 1.1 Wh 0.42 gCO2e
96.75B* US
Documented region, country retained as reference
6.4 Wh 2.5 gCO2e 0.19 Wh 0.0714 gCO2e
96.75B* US
Documented region, country retained as reference
6.7 Wh 2.6 gCO2e 0.19 Wh 0.0748 gCO2e
600B* US
Documented region, country retained as reference
36.3 Wh 14.0 gCO2e 1.0 Wh 0.40 gCO2e
600B* US
Documented region, country retained as reference
38.0 Wh 14.6 gCO2e 1.1 Wh 0.42 gCO2e
4000B* US
Documented region, country retained as reference
220.2 Wh 84.8 gCO2e 6.4 Wh 2.5 gCO2e
178B US
Documented region, country retained as reference
9.3 Wh 3.6 gCO2e 0.27 Wh 0.10 gCO2e
137B* US
Documented region, country retained as reference
7.3 Wh 2.8 gCO2e 0.21 Wh 0.0807 gCO2e
405B US
Comparative reference country
19.1 Wh 7.4 gCO2e 0.55 Wh 0.21 gCO2e
70B US
Comparative reference country
3.6 Wh 1.4 gCO2e 0.10 Wh 0.0402 gCO2e
8B US
Comparative reference country
0.46 Wh 0.18 gCO2e 0.0133 Wh 0.0051 gCO2e
17B active / 400B total US
Comparative reference country
1.4 Wh 0.55 gCO2e 0.0410 Wh 0.0158 gCO2e
17B active / 109B total US
Comparative reference country
1.4 Wh 0.54 gCO2e 0.0402 Wh 0.0155 gCO2e
530B US
Screening proxy
26.2 Wh 10.1 gCO2e 0.76 Wh 0.29 gCO2e
14B FR
Provider-country proxy
0.88 Wh 0.0346 gCO2e 0.0255 Wh 0.0010 gCO2e
3B FR
Provider-country proxy
0.19 Wh 0.0069 gCO2e 0.00560 Wh 0.0002 gCO2e
8B FR
Provider-country proxy
0.49 Wh 0.0208 gCO2e 0.0142 Wh 0.0006 gCO2e
123B FR
Provider-country proxy
7.2 Wh 0.29 gCO2e 0.21 Wh 0.0083 gCO2e
41B active / 675B total FR
Provider-country proxy
3.2 Wh 0.13 gCO2e 0.0925 Wh 0.0037 gCO2e
128B* FR
Provider-country proxy
7.2 Wh 0.29 gCO2e 0.21 Wh 0.0084 gCO2e
24B FR
Provider-country proxy
1.4 Wh 0.0588 gCO2e 0.0413 Wh 0.0017 gCO2e
6.5B active / 119B total FR
Provider-country proxy
0.56 Wh 0.0208 gCO2e 0.0162 Wh 0.0006 gCO2e
320B* US
Screening proxy
18.2 Wh 7.0 gCO2e 0.52 Wh 0.20 gCO2e
200B* US
Screening proxy
11.6 Wh 4.5 gCO2e 0.34 Wh 0.13 gCO2e
320B* US
Screening proxy
18.1 Wh 7.0 gCO2e 0.52 Wh 0.20 gCO2e
175B US
Documented region, country retained as reference
8.1 Wh 3.1 gCO2e 0.23 Wh 0.0900 gCO2e
32B CN
Provider-country proxy
1.7 Wh 0.93 gCO2e 0.0496 Wh 0.0268 gCO2e
72B CN
Provider-country proxy
3.7 Wh 2.0 gCO2e 0.11 Wh 0.0579 gCO2e
7B CN
Provider-country proxy
0.40 Wh 0.22 gCO2e 0.0117 Wh 0.0063 gCO2e
22B active / 235B total CN
Provider-country proxy
1.5 Wh 0.82 gCO2e 0.0437 Wh 0.0236 gCO2e
32.8B CN
Provider-country proxy
1.8 Wh 0.95 gCO2e 0.0508 Wh 0.0274 gCO2e

`Retained country` is the country actually used to recalculate CO2 via the electricity mix. When the exact country is not published, the project uses an explicit screening proxy rather than presenting a location as certain.

`*` indicates an estimated parameter count rather than a provider-published value.

The market-model comparison now relies on market_multifactor_prompt_proxy_v1: a prompt-energy screening proxy whose main prompt-level calibration anchor comes from Elsworth et al. (2025), then adjusted by active parameters, context window, serving mode, modality support, architecture overhead, and standardized token volume, and interpreted alongside other inference references.

Models

83 market models with quantified training impacts

This table projects the training orders of magnitude of the quantified market models from the indicator families actually available in the literature: training energy derived from emissions when the source country is documented in the electricity-mix table, and training CO2e contextualized from that retained energy and the model country proxy. The current screening proxy combines retained parameter count, a training-token prior, a training-regime prior, architecture features, and a hardware-class proxy. The quantified layer combines 56 strict rows and 27 partial-data donor priors.

100B* 186.2 MWh 71.71 tCO2e
100B* 41.0 MWh 15.78 tCO2e
100B* 484.2 MWh 186.4 tCO2e
100B* 223.5 MWh 86.05 tCO2e
22B* 9.0 MWh 3.47 tCO2e
175B* 570.4 MWh 219.6 tCO2e
175B* 570.4 MWh 219.6 tCO2e
5000B* 50.9 GWh 19 614 tCO2e
2000B* 74.5 GWh 28 682 tCO2e
5000B* 98.0 GWh 37 719 tCO2e
4000B* 62.7 GWh 24 140 tCO2e
175B* 1.3 GWh 501.9 tCO2e
400B* 3.0 GWh 1 147 tCO2e
22B 8.7 MWh 0.35 tCO2e
37B active / 671B total 7.3 GWh 3 938 tCO2e
37B active / 671B total 7.3 GWh 3 938 tCO2e
13B active / 284B total 1.3 GWh 705.4 tCO2e
49B active / 1600B total 41.5 GWh 22 389 tCO2e
123B 272.2 MWh 10.89 tCO2e
133.5B* 149.2 MWh 57.44 tCO2e
133.5B* 298.4 MWh 114.9 tCO2e
30B* 16.8 MWh 6.45 tCO2e
80B* 119.2 MWh 45.89 tCO2e
30B* 16.8 MWh 6.45 tCO2e
200B* 745.0 MWh 286.8 tCO2e
30B* 55.9 MWh 21.51 tCO2e
3000B* 185.7 GWh 71 492 tCO2e
30B* 55.9 MWh 21.51 tCO2e
22B* 16.4 MWh 6.31 tCO2e
22B* 16.4 MWh 6.31 tCO2e
3000B* 185.7 GWh 71 492 tCO2e
130B 379.0 MWh 145.9 tCO2e
280B 1.4 GWh 244.9 tCO2e
175B* 578.6 MWh 222.8 tCO2e
440B* active / 1760B* total 8.9 GWh 3 407 tCO2e
200B* 745.0 MWh 286.8 tCO2e
200B* 29.8 MWh 11.47 tCO2e
95B* 168.1 MWh 64.71 tCO2e
95B* 168.1 MWh 64.71 tCO2e
3000B* 185.7 GWh 71 492 tCO2e
440B* 2.5 GWh 946.5 tCO2e
3000B* 185.7 GWh 71 492 tCO2e
95B* 168.1 MWh 64.71 tCO2e
95B* 168.1 MWh 64.71 tCO2e
3000B* 111.4 GWh 42 895 tCO2e
3000B* 123.8 GWh 47 661 tCO2e
3000B* 123.8 GWh 47 661 tCO2e
5.1B active / 117B total 232.8 MWh 89.62 tCO2e
3.6B active / 21B total 7.5 MWh 2.89 tCO2e
78.5B* active / 314B* total 1.7 GWh 636.3 tCO2e
78.5B* 114.4 MWh 44.05 tCO2e
115B* active / 270B* total 1.2 GWh 470.5 tCO2e
600B* 6.7 GWh 2 581 tCO2e
96.75B* 216.2 MWh 83.25 tCO2e
96.75B* 216.2 MWh 83.25 tCO2e
600B* 7.3 GWh 2 797 tCO2e
600B* 7.3 GWh 2 797 tCO2e
4000B* 10.2 GWh 3 923 tCO2e
178B 840.8 MWh 323.7 tCO2e
137B* 332.8 MWh 128.1 tCO2e
405B 3.1 GWh 1 193 tCO2e
70B 92.6 MWh 35.64 tCO2e
8B 1.1 MWh 0.42 tCO2e
17B active / 400B total 8.6 GWh 3 313 tCO2e
17B active / 109B total 4.3 GWh 1 641 tCO2e
530B 8.6 GWh 3 305 tCO2e
14B 4.5 MWh 0.18 tCO2e
3B 179.4 kWh 0.01 tCO2e
8B 1.0 MWh 0.04 tCO2e
123B 281.8 MWh 11.27 tCO2e
41B active / 675B total 8.5 GWh 339.4 tCO2e
128B* 339.1 MWh 13.56 tCO2e
24B 11.9 MWh 0.48 tCO2e
6.5B active / 119B total 263.7 MWh 10.55 tCO2e
320B* 1.9 GWh 734.3 tCO2e
200B* 307.7 MWh 118.5 tCO2e
320B* 1.6 GWh 598.6 tCO2e
175B 661.3 MWh 254.6 tCO2e
32B 19.3 MWh 10.45 tCO2e
72B 97.9 MWh 52.89 tCO2e
7B 826.0 kWh 0.45 tCO2e
22B active / 235B total 939.1 MWh 507.1 tCO2e
32.8B 20.3 MWh 10.98 tCO2e

`*` indicates an estimated parameter count rather than a provider-published value.

ImpactLLM is designed as a transparent screening tool, not as a black-box score. The current release starts from source-linked inference anchors, then exposes a bounded multi-factor proxy rather than a hidden single-number score.

1. Source-linked literature anchors.

The application-level estimator starts from published inference indicators linked to an explicit source, model, geography, and system boundary. In the current market-model release, the predictive core uses Elsworth et al. (2025) as the main prompt-level calibration anchor, with a median prompt energy of 0.24 Wh/prompt for Gemini Apps, and is interpreted alongside other inference references such as the ML.ENERGY Benchmark, Ren et al. (2024), and Li et al. (2025).

2. A multi-factor effective-parameter proxy.

When direct telemetry is unavailable for a target model, ImpactLLM does not rely on a raw parameter multiple alone. It builds an effective active-parameter profile from the retained model characteristics: active parameters, context window, serving mode (open, hybrid, closed), modality support, and architecture notes such as MoE or reasoning-oriented overheads.

3. Token volume remains explicit.

The current proxy adjusts the anchor with a weighted prompt-compute volume defined from input and output tokens. Output generation is weighted more heavily than input processing, so output-heavy scenarios and repeated LLM calls raise the estimate materially.

The current prompt-level branch is a screening proxy, not an audited benchmark. For this reason, the application returns a bounded low-central-high result rather than one falsely precise deterministic value.

4. Carbon derived from context.

Carbon is not copied mechanically from the source paper. It is recalculated from the retained energy estimate using the electricity mix associated with the selected country context.

5. A research-oriented estimator.

The result is an auditable estimate intended for comparison, software design, and methodological discussion. It is useful precisely because the assumptions, factors, and retained sources remain visible and inspectable.

Technical paper

Pachot, A., & Petit, T. (2026, March 14). Transparent Screening for LLM Inference and Training Impacts. /impact-llm/downloads/ImpactLLM_paper.pdf

We work on responsible AI with a focus on methodological rigor, traceability, and real-world decision support. Our work combines scientific research, product design, and operational deployment to make AI systems more transparent, more accountable, and more useful in practice.

How to cite ImpactLLM

Pachot, A., & Petit, T. (2026, March 14). Transparent Screening for LLM Inference and Training Impacts.

BibTeX

@misc{impactllm_screening_2026,
  title = {Transparent Screening for LLM Inference and Training Impacts},
  author = {Pachot, Arnault and Petit, Thierry},
  year = {2026},
  month = mar,
  note = {Conference paper preprint},
  url = {https://dev.emotia.com/impact-llm/downloads/ImpactLLM_paper.pdf}
}

Download technical paper PDF | Download technical paper BibTeX

GitHub repository

The project repository is available on GitHub: https://github.com/apachot/ImpactLLM.

License

This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.

This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.

You should have received a copy of the GNU General Public License along with this program. If not, see https://www.gnu.org/licenses/.

Arnault Pachot

Arnault Pachot is a researcher and entrepreneur, founder of OpenStudio and now founder of Emotia. He works on responsible digital transformation, Green IT, and decision-oriented AI systems. He co-authored the Dunod book Intelligence artificielle et environnement : alliance ou nuisance ?, dedicated to practical pathways for environmentally responsible AI.

Emotia

LinkedIn: Arnault Pachot

Google Scholar: Arnault Pachot

Thierry Petit

Thierry Petit is a senior AI researcher and scientific leader with more than twenty years of academic and R&D experience in Europe and the United States. His work spans trustworthy AI, simulation, optimization, and decision-grade platforms. At Emotia and Pollitics, he leads the scientific direction of systems designed to remain both operationally useful and methodologically robust.

Emotia

LinkedIn: Thierry Petit

Google Scholar: Thierry Petit

Selected references on AI and the environment

Sources

This annex brings together the quantified reference material used in the interface, along with everyday comparison benchmarks and country factors used for carbon and water recalculation.

Inference reference set

Ref. Data type LLM model Parameters Country Value Citation
[43]Energy consumption per prompt (Gemini Apps median prompt)Gemini Apps (version non spécifiée)180BNon spécifié0.24 Wh/promptElsworth, C., Huang, K., Patterson, D., Schneider, I., & others (2025). Measuring the Environmental Impact of Delivering AI at Google Scale. arXiv preprint arXiv:2508.15734. https://arxiv.org/abs/2508.15734
Table 2, p. 7
[44]Emissions per prompt (Gemini Apps median prompt)Gemini Apps (version non spécifiée)180BNon spécifié0.03 gCO2e/promptElsworth, C., Huang, K., Patterson, D., Schneider, I., & others (2025). Measuring the Environmental Impact of Delivering AI at Google Scale. arXiv preprint arXiv:2508.15734. https://arxiv.org/abs/2508.15734
Table 2, p. 7
[46]Energy consumption to generate one page (Llama-3-70B, 500-word page)Llama 3 70B70BÉtats-Unis0.0195 kWh/pageRen, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the Contrasting Narratives on the Environmental Impact of Large Language Models. Scientific Reports, 14, 28180.
Results section, Llama-3-70B paragraphs
[47]Emissions to generate one page (Llama-3-70B, 500-word page)Llama 3 70B70BÉtats-Unis15 gCO2/pageRen, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the Contrasting Narratives on the Environmental Impact of Large Language Models. Scientific Reports, 14, 28180.
Results section, Llama-3-70B paragraphs
[49]Energy consumption to generate one page (Gemma-2B-it, 500-word page)Gemma-2B-it2BÉtats-Unis0.00024 kWh/pageRen, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the Contrasting Narratives on the Environmental Impact of Large Language Models. Scientific Reports, 14, 28180.
Results section, Gemma-2B-it paragraphs
[50]Emissions to generate one page (Gemma-2B-it, 500-word page)Gemma-2B-it2BÉtats-Unis0.18 gCO2/pageRen, S., Tomlinson, B., Black, R. W., & Torrance, A. W. (2024). Reconciling the Contrasting Narratives on the Environmental Impact of Large Language Models. Scientific Reports, 14, 28180.
Results section, Gemma-2B-it paragraphs

Training reference set

Ref. Data type LLM model Parameters Country Value Citation
[1]Greenhouse gas emissions from training (Transformer (big))Transformer (big)213MÉtats-Unis192 lb CO2eStrubell, E., Ganesh, A., & McCallum, A. (2019). Energy and Policy Considerations for Deep Learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645--3650.
Table 1, p. 1
[2]Greenhouse gas emissions from training (BLOOM 176B)BLOOM 176B176BFrance24.7 tCO2eLuccioni, A. S., Viguier, S., & Ligozat, A. L. (2023). Estimating the Carbon Footprint of BLOOM. Journal of Machine Learning Research, 24(253), 1--15. https://www.jmlr.org/papers/v24/23-0069.html
Abstract, p. 1
[3]Greenhouse gas emissions from training (BLOOM 176B)BLOOM 176B176BFrance50.5 tCO2eLuccioni, A. S., Viguier, S., & Ligozat, A. L. (2023). Estimating the Carbon Footprint of BLOOM. Journal of Machine Learning Research, 24(253), 1--15. https://www.jmlr.org/papers/v24/23-0069.html
Abstract, p. 1; Table 3, p. 7
[4]Greenhouse gas emissions from training (Llama 3.1 8B)Llama 3.1 8B8BNon spécifié420 tCO2eMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions'
[5]Greenhouse gas emissions from training (Llama 3.1 70B)Llama 3.1 70B70BNon spécifié2040 tCO2eMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions'
[6]Greenhouse gas emissions from training (Llama 3.1 405B)Llama 3.1 405B405BNon spécifié8930 tCO2eMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions'
[7]Compute time used for training (Llama 3.1 405B)Llama 3.1 405B405BNon spécifié30.84 million GPU-hoursMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', cumulative compute table
[8]Compute time used for training (Llama 3.1 8B)Llama 3.1 8B8BNon spécifié1.46 million GPU-hoursMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', cumulative compute table
[9]Compute time used for training (Llama 3.1 70B)Llama 3.1 70B70BNon spécifié7.0 million GPU-hoursMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Hardware and Software', cumulative compute table
[10]Training token volume (Llama 3.1 8B)Llama 3.1 8B8BNon spécifié15 trillion tokensMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Training Data'
[11]Training token volume (Llama 3.1 70B)Llama 3.1 70B70BNon spécifié15 trillion tokensMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Training Data'
[12]Training token volume (Llama 3.1 405B)Llama 3.1 405B405BNon spécifié15 trillion tokensMeta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B
Section 'Training Data'
[13]Greenhouse gas emissions from training (Llama 4 Scout)Llama 4 Scout109BNon spécifié1354 tCO2eMeta (2025) Llama 4 Model Card
Section 'Hardware and Software', training emissions table
[14]Greenhouse gas emissions from training (Llama 4 Maverick)Llama 4 Maverick400BNon spécifié645 tCO2eMeta (2025) Llama 4 Model Card
Section 'Hardware and Software', training emissions table
[15]Training : training energy (Llama 4 Scout)Llama 4 Scout109BNon spécifié3.5 GWhMeta (2025) Llama 4 Model Card
Section 'Hardware and Software', 5.0M GPU hours on H100-80GB at 700W
[16]Training : training energy (Llama 4 Maverick)Llama 4 Maverick400BNon spécifié1.666 GWhMeta (2025) Llama 4 Model Card
Section 'Hardware and Software', 2.38M GPU hours on H100-80GB at 700W
[17]Training token volume (Llama 4 Scout)Llama 4 Scout109BNon spécifié40 trillion tokensMeta (2025) Llama 4 Model Card
Section 'Training Data'
[18]Training token volume (Llama 4 Maverick)Llama 4 Maverick400BNon spécifié22 trillion tokensMeta (2025) Llama 4 Model Card
Section 'Training Data'
[19]Emissions across the model creation lifecycle (OLMo 20M)OLMo 20M20MÉtats-Unis0.3 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[20]Emissions across the model creation lifecycle (OLMo 60M)OLMo 60M60MÉtats-Unis0.4 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[21]Emissions across the model creation lifecycle (OLMo 150M)OLMo 150M150MÉtats-Unis1 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[22]Emissions across the model creation lifecycle (OLMo 300M)OLMo 300M300MÉtats-Unis2 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[23]Emissions across the model creation lifecycle (OLMo 700M)OLMo 700M700MÉtats-Unis3 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[24]Emissions across the model creation lifecycle (OLMo 7B)OLMo 7B7BÉtats-Unis22 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[25]Emissions across the model creation lifecycle (OLMo 1B (3T))OLMo 1B (3T)1BÉtats-Unis10 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[26]Emissions across the model creation lifecycle (OLMo 7B (Twin))OLMo 7B (Twin)7BÉtats-Unis70 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[27]Emissions across the model creation lifecycle (OLMo (04|07)24 7B)OLMo (04|07)24 7B7BÉtats-Unis32 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[28]Emissions across the model creation lifecycle (OLMo 2 7B)OLMo 2 7B7BÉtats-Unis52 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[29]Emissions across the model creation lifecycle (OLMo 2 13B)OLMo 2 13B13BÉtats-Unis101 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4
[30]Emissions across the model creation lifecycle (OLMoE 0924)OLMoE 09241B active / 7B totalÉtats-Unis18 tCO2eMorrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804
Table 2, p. 6; cluster locations in Section 3.1, p. 4

Real-world comparison benchmarks

No. Domain Indicator Value Reference
1 Energy Fluorescent lamp for 1 hour 9,3 Wh Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf
Purdue Extension guide used for household-use examples; extraction still to be normalized more precisely.
2 Energy Laptop for 1 hour 32 Wh Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf
Purdue Extension guide used for household-use examples; extraction still to be normalized more precisely.
3 Energy Laptop for 3 hours 96 Wh Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf
Project derivation from the Purdue laptop example retained at 32 Wh for 1 hour, rescaled here to 3 hours of use.
4 Energy Electric space heater for 1 hour 1.5 kWh Project calculation convention.
Nominal power assumption fixed at 1,500 W, i.e. 1.5 kWh for 1 hour.
5 Energy Laptop for 5 hours 160 Wh Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf
Project derivation from the Purdue laptop example retained at 32 Wh for 1 hour, rescaled here to 5 hours of use.
6 Energy Electric kettle for 7 minutes 210 Wh Project calculation convention.
Nominal power assumption fixed at 1,800 W, i.e. 210 Wh for 7 minutes of use.
7 Energy Electric space heater for 10 minutes 250 Wh Project calculation convention.
Project derivation from the retained 1,500 W space-heater convention, rescaled to 10 minutes of use.
8 Energy 10,000 French households over one year of domestic use 25 GWh RTE. (2022, February 25). Bilan électrique 2021 - Une production d’électricité assurée à plus de 92% par des sources n’émettant pas de gaz à effet de serre. https://www.rte-france.com/actualites/bilan-electrique-2021
Project convention based on an average consumption of 2,500 kWh/year per household; comparison value made explicit in the training-chart note.
9 Carbon Average gasoline car for 1 km 235 gCO2e International Council on Clean Transportation. (2025). Life-cycle greenhouse gas emissions from passenger cars in the European Union: A 2025 update and key factors to consider. https://theicct.org/publication/electric-cars-life-cycle-analysis-emissions-europe-jul25/
Key findings ; Figure 1 ; gasoline ICEV running on the average blend of fossil gasoline and ethanol estimated at 235 gCO2e/km, i.e. 235 gCO2e over 1 km.
10 Carbon Average gasoline car for 430 m 101.1 gCO2e International Council on Clean Transportation. (2025). Life-cycle greenhouse gas emissions from passenger cars in the European Union: A 2025 update and key factors to consider. https://theicct.org/publication/electric-cars-life-cycle-analysis-emissions-europe-jul25/
Project derivation from the ICCT benchmark of 235 gCO2e/km for an average gasoline car, rescaled here to 430 meters to align with the current upper-end inference carbon tier retained in the comparison.
11 Carbon Average full commercial flight (derived value) ≈ 21.1 tCO2 per flight Gössling, S., Klöwer, M., Leitão, J. C., Hirsch, S., Brockhagen, D., & Humpe, A. (2026). Large carbon dioxide emissions avoidance potential in improved commercial air transport efficiency. Communications Earth & Environment, 7, 13. https://www.nature.com/articles/s43247-025-03069-4
Results, “Emissions and efficiency”: 27,451,887 flights in 2023 causing 577,968,750 tCO2 emissions; the site then derives an average per flight.

Central screening factors retained for market models

This table documents the central values retained by the project for the multi-factor prompt proxy of each quantified market model: raw active parameters, context window, serving mode, modality support, the resulting central factors F_ctx, F_srv, F_mod, F_arch, and the resulting central effective active-parameter proxy P_eff,c. The current annex covers 56 strict rows and 27 partial-data rows derived from donor models with sourced parameter counts.

Model Provider Active parameters Context window Serving mode Vision F_ctx F_srv F_mod F_arch P_eff,c
ai21 178B [1] 2,048 [2] closed [4] no [3] 1.000 1.140 1.000 1.000 202.920B
alibaba 32B [7] 131,072 [8] open [10] no [9] 1.070 1.000 1.000 1.000 34.240B
alibaba 72B [13] 131,072 [14] open [16] no [15] 1.070 1.000 1.000 1.000 77.040B
alibaba 7B [19] 131,072 [20] open [22] no [21] 1.070 1.000 1.000 1.000 7.490B
alibaba 22B active / 235B total [25] 131,072 [26] open [28] no [27] 1.070 1.000 1.000 1.273 29.975B
alibaba 32.8B [31] 131,072 [32] open [34] no [33] 1.070 1.000 1.000 1.000 35.096B
anthropic 100B* [37] 200,000 [38] closed [40] yes [39] 1.091 1.140 1.030 1.000 128.145B
anthropic 100B* [43] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 128.145B
anthropic 100B* [43] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 128.145B
anthropic 100B* [43] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 128.145B
anthropic 22B* [48] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 28.192B
anthropic 175B* [37] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 224.253B
anthropic 175B* [37] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 224.253B
anthropic 5000B* [43] 1,000,000 [44] research [46] yes [45] 1.173 1.140 1.030 1.000 6884.363B
anthropic 2000B* [37] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 2562.897B
anthropic 5000B* [37] 1,000,000 [44] closed [46] yes [45] 1.173 1.140 1.030 1.000 6884.363B
anthropic 4000B* [50] 1,000,000 [44] closed [46] yes [45] 1.173 1.140 1.030 1.000 5507.491B
anthropic 175B* [43] 200,000 [44] closed [46] yes [45] 1.091 1.140 1.030 1.000 224.253B
anthropic 400B* [37] 1,000,000 [44] closed [46] yes [45] 1.173 1.140 1.030 1.000 550.749B
deepmind 280B [52] 4,096 [53] closed [55] no [54] 1.000 1.140 1.000 1.000 319.200B
deepseek 37B active / 671B total [58] 128,000 [59] hybrid [61] no [60] 1.069 1.070 1.000 1.401 59.289B
deepseek 37B active / 671B total [64] 128,000 [65] hybrid [67] no [66] 1.069 1.070 1.000 1.334 56.466B
deepseek 13B active / 284B total [70] 1,000,000 [71] hybrid [73] no [72] 1.173 1.070 1.000 1.356 22.117B
deepseek 49B active / 1600B total [76] 1,000,000 [77] hybrid [79] no [78] 1.173 1.070 1.000 1.472 90.526B
google 133.5B* [43] 1,048,576 [82] closed [84] yes [83] 1.175 1.140 1.030 1.000 184.188B
google 133.5B* [43] 1,048,576 [86] closed [88] yes [87] 1.175 1.140 1.030 1.000 184.188B
google 30B* [37] 1,048,576 [90] closed [92] yes [91] 1.175 1.140 1.030 1.000 41.391B
google 80B* [94] 1,048,576 [90] closed [92] yes [91] 1.175 1.140 1.030 1.000 110.375B
google 30B* [94] 1,048,576 [90] closed [92] yes [91] 1.175 1.140 1.030 1.000 41.391B
google 200B* [37] 1,048,576 [90] closed [92] yes [91] 1.175 1.140 1.030 1.000 275.937B
google 30B* [43] 1,048,576 [96] closed [98] yes [97] 1.175 1.140 1.030 1.000 41.391B
google 3000B* [37] 1,048,576 [100] closed [102] yes [101] 1.175 1.140 1.030 1.000 4139.055B
google 30B* [43] 131,072 [104] closed [106] yes [105] 1.070 1.140 1.030 1.000 37.692B
google 22B* [43] 8,192 [108] closed [110] no [109] 1.000 1.140 1.000 1.000 25.080B
google 22B* [43] 1,048,576 [112] closed [114] yes [113] 1.175 1.140 1.030 1.000 30.353B
google 3000B* [37] 1,048,576 [116] closed [118] yes [117] 1.175 1.140 1.030 1.000 4139.055B
google 130B [120] 2,048 [121] research [123] no [122] 1.000 1.140 1.000 1.000 148.200B
google 137B* [126] 2,048 [127] closed [129] no [128] 1.000 1.140 1.000 1.000 156.180B
meta 405B [132] 131,072 [133] open [135] no [134] 1.070 1.000 1.000 1.000 433.350B
meta 70B [132] 131,072 [138] open [140] no [139] 1.070 1.000 1.000 1.000 74.900B
meta 8B [132] 131,072 [142] open [144] no [143] 1.070 1.000 1.000 1.000 8.560B
meta 17B active / 400B total [146] 1,000,000 [147] open [149] yes [148] 1.173 1.000 1.030 1.365 28.017B
meta 17B active / 109B total [153] 10,000,000 [154] open [156] yes [155] 1.289 1.000 1.030 1.214 27.408B
meta 175B [160] 2,048 [161] open [163] no [162] 1.000 1.000 1.000 1.000 175.000B
microsoft 530B [166] 2,048 [167] closed [169] no [168] 1.000 1.140 1.000 1.000 604.200B
mistral 22B [172] 32,000 [173] hybrid [175] no [174] 1.000 1.070 1.000 1.000 23.540B
mistral 123B [178] 256,000 [179] hybrid [181] no [180] 1.104 1.070 1.000 1.000 145.271B
mistral 14B [184] 256,000 [185] hybrid [187] yes [186] 1.104 1.070 1.030 1.000 17.031B
mistral 3B [172] 128,000 [190] hybrid [192] no [191] 1.069 1.070 1.000 1.000 3.431B
mistral 8B [172] 128,000 [194] hybrid [196] no [195] 1.069 1.070 1.000 1.000 9.149B
mistral 123B [198] 128,000 [199] closed [201] yes [200] 1.069 1.140 1.030 1.000 154.364B
mistral 41B active / 675B total [204] 256,000 [205] hybrid [207] yes [206] 1.104 1.070 1.030 1.323 66.001B
mistral 128B* [210] 256,000 [211] hybrid [213] yes [212] 1.104 1.070 1.030 1.000 155.712B
mistral 24B [172] 128,000 [216] hybrid [218] yes [217] 1.069 1.070 1.030 1.000 28.270B
mistral 6.5B active / 119B total [220] 256,000 [221] hybrid [223] yes [222] 1.104 1.070 1.030 1.336 10.561B
openai 175B* [226] 4,096 [227] closed [229] no [228] 1.000 1.140 1.000 1.000 199.500B
openai 440B* active / 1760B* total [233] 8,192 [234] closed [236] yes [235] 1.000 1.140 1.030 1.160 599.312B
openai 200B* [50] 128,000 [241] closed [243] yes [242] 1.069 1.140 1.030 1.000 250.998B
openai 200B* [43] 128,000 [245] closed [247] yes [246] 1.069 1.140 1.030 1.000 250.998B
openai 95B* [249] 400,000 [250] closed [252] yes [251] 1.126 1.140 1.030 1.000 125.642B
openai 95B* [255] 400,000 [256] closed [258] yes [257] 1.126 1.140 1.030 1.000 125.642B
openai 3000B* [37] 400,000 [261] closed [263] yes [262] 1.126 1.140 1.030 1.000 3967.636B
openai 440B* [43] 400,000 [265] closed [267] yes [266] 1.126 1.140 1.030 1.000 581.920B
openai 3000B* [37] 400,000 [269] closed [271] yes [270] 1.126 1.140 1.030 1.000 3967.636B
openai 95B* [43] 400,000 [273] closed [275] yes [274] 1.126 1.140 1.030 1.000 125.642B
openai 95B* [43] 400,000 [277] closed [279] yes [278] 1.126 1.140 1.030 1.000 125.642B
openai 3000B* [43] 400,000 [281] closed [283] yes [282] 1.126 1.140 1.030 1.000 3967.636B
openai 3000B* [43] 400,000 [285] closed [287] yes [286] 1.126 1.140 1.030 1.000 3967.636B
openai 3000B* [43] 1,050,000 [289] closed [291] yes [290] 1.175 1.140 1.030 1.000 4139.296B
openai 5.1B active / 117B total [293] 131,072 [294] open [296] no [295] 1.070 1.000 1.000 1.362 7.430B
openai 3.6B active / 21B total [299] 131,072 [300] open [302] no [301] 1.070 1.000 1.000 1.204 4.636B
openai 320B* [43] 200,000 [305] closed [307] yes [306] 1.091 1.140 1.030 1.000 410.063B
openai 200B* [43] 128,000 [309] closed [311] no [310] 1.069 1.140 1.000 1.050 255.871B
openai 320B* [43] 128,000 [313] closed [315] no [314] 1.069 1.140 1.000 1.050 409.394B
xai 78.5B* active / 314B* total [317] 200,000 [318] closed [320] yes [319] 1.091 1.140 1.030 1.160 116.689B
xai 78.5B* [43] 128,000 [324] closed [326] no [325] 1.069 1.140 1.000 1.050 100.429B
xai 115B* active / 270B* total [328] 2,000,000 [318] closed [320] yes [319] 1.208 1.140 1.030 1.099 179.130B
xai 600B* [37] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.050 893.321B
xai 96.75B* [43] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.000 137.189B
xai 96.75B* [43] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.050 144.048B
xai 600B* [43] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.000 850.782B
xai 600B* [43] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.050 893.321B
xai 4000B* [43] 2,000,000 [324] closed [326] yes [325] 1.208 1.140 1.030 1.000 5671.879B

Central training screening factors retained for market models

This table documents the central values retained by the project for the multi-factor training proxy of each quantified market model: retained training parameter count, training-token prior, training regime, multimodal training flag, hardware-class proxy, and the resulting central factors F_reg, F_arch-tr, and F_hw. These are project screening factors, not provider-published measurements. When a field is retained as an estimate or screening prior, its citation anchors the release line or methodological basis rather than an exact provider-published numeric value.

Model Provider Retained parameters Training tokens Training regime Multimodal Hardware class F_reg F_arch F_hw
ai21 178B [5] 5.00T* pretraining* no [6] standard_gpu_cluster 1.0000 1.0000 1.0500
alibaba 32B [11] 0.64T* pretraining* no [12] standard_gpu_cluster* 1.0000 1.0000 1.0500
alibaba 72B [17] 1.44T* pretraining* no [18] standard_gpu_cluster* 1.0000 1.0000 1.0500
alibaba 7B [23] 0.14T* pretraining* no [24] standard_gpu_cluster* 1.0000 1.0000 1.0500
alibaba 235B [29] 4.70T* pretraining* no [30] standard_gpu_cluster* 1.0000 0.9000 1.0500
alibaba 32.8B [35] 0.66T* pretraining* no [36] standard_gpu_cluster* 1.0000 1.0000 1.0500
anthropic 100B [41] 2.00T* pretraining* yes [42] modern_hyperscale_gpu 1.0000 1.1500 0.9000
anthropic 100B* 0.44T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 100B* 5.20T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 100B* 2.40T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 22B [49] 0.44T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 175B [41] 3.50T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 175B [41] 3.50T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 5000B* 52.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 2000B [41] 40.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 5000B [41] 100.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 4000B [51] 80.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 175B* 8.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
anthropic 400B [41] 8.00T* pretraining* yes [47] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
deepmind 280B [56] 6.00T* pretraining* no [57] modern_hyperscale_gpu 1.0000 1.0000 0.9000
deepseek 671B [62] 13.42T* pretraining* no [63] mixed_gpu_cluster* 1.0000 0.9000 1.0000
deepseek 671B [68] 13.42T* pretraining* no [69] mixed_gpu_cluster* 1.0000 0.9000 1.0000
deepseek 284B [74] 5.68T* pretraining* no [75] mixed_gpu_cluster* 1.0000 0.9000 1.0000
deepseek 1600B [80] 32.00T* pretraining* no [81] mixed_gpu_cluster* 1.0000 0.9000 1.0000
google 133.5B* 1.20T* pretraining* yes [85] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 133.5B* 2.40T* pretraining* yes [89] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 30B [41] 0.60T* pretraining* yes [93] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 80B [95] 1.60T* pretraining* yes [93] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 30B [95] 0.60T* pretraining* yes [93] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 200B [41] 4.00T* pretraining* yes [93] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 30B* 2.00T* pretraining* yes [99] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 3000B [41] 60.00T* pretraining* yes [103] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 30B* 2.00T* pretraining* yes [107] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 22B* 0.80T* pretraining* yes [111] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 22B* 0.80T* pretraining* yes [115] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 3000B [41] 60.00T* pretraining* yes [119] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
google 130B [124] 4.00T* pretraining* no [125] modern_hyperscale_gpu 1.0000 0.9000 0.9000
google 137B [130] 3.00T* pretraining* no [131] modern_hyperscale_gpu 1.0000 1.0000 0.9000
meta 405B [136] 8.10T* pretraining* no [137] standard_gpu_cluster* 1.0000 1.0000 1.0500
meta 70B [136] 1.40T* pretraining* no [141] standard_gpu_cluster* 1.0000 1.0000 1.0500
meta 8B [136] 0.16T* pretraining* no [145] standard_gpu_cluster* 1.0000 1.0000 1.0500
meta 400B [150] 22.00T [151] pretraining* yes [152] standard_gpu_cluster* 1.0000 1.0350 1.0500
meta 109B [157] 40.00T [158] pretraining* yes [159] standard_gpu_cluster* 1.0000 1.0350 1.0500
meta 175B [164] 4.00T* pretraining* no [165] standard_gpu_cluster 1.0000 1.0000 1.0500
microsoft 530B [170] 20.00T* pretraining* no [171] modern_hyperscale_gpu 1.0000 1.0000 0.9000
mistral 22B [176] 0.44T* pretraining* no [177] mixed_gpu_cluster* 1.0000 1.0000 1.0000
mistral 123B [182] 2.46T* pretraining* no [183] mixed_gpu_cluster* 1.0000 1.0000 1.0000
mistral 14B [188] 0.28T* pretraining* yes [189] mixed_gpu_cluster* 1.0000 1.1500 1.0000
mistral 3B [176] 0.06T* pretraining* no [193] mixed_gpu_cluster* 1.0000 1.0000 1.0000
mistral 8B [176] 0.16T* pretraining* no [197] mixed_gpu_cluster* 1.0000 1.0000 1.0000
mistral 123B [202] 2.46T* pretraining* yes [203] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
mistral 675B [208] 13.50T* pretraining* yes [209] mixed_gpu_cluster* 1.0000 1.0350 1.0000
mistral 128B [214] 2.56T* pretraining* yes [215] mixed_gpu_cluster* 1.0000 1.1500 1.0000
mistral 24B [176] 0.48T* pretraining* yes [219] mixed_gpu_cluster* 1.0000 1.1500 1.0000
mistral 119B [224] 2.38T* pretraining* yes [225] mixed_gpu_cluster* 1.0000 1.0350 1.0000
openai 175B [230] 3.50T* pretraining* no [231] standard_gpu_cluster [232] 1.0000 1.0000 1.0500
openai 1760B [237] 6.00T* pretraining [238] yes [239] modern_hyperscale_gpu [240] 1.0000 1.0350 0.9000
openai 200B [51] 4.00T* pretraining* yes [244] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 200B* 0.16T* pretraining* yes [248] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 95B [253] 1.90T* pretraining* yes [254] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 95B [259] 1.90T* pretraining* yes [260] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 3000B [41] 60.00T* pretraining* yes [264] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 440B* 6.00T* pretraining* yes [268] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 3000B [41] 60.00T* pretraining* yes [272] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 95B* 1.90T* pretraining* yes [276] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 95B* 1.90T* pretraining* yes [280] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 3000B* 36.00T* pretraining* yes [284] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 3000B* 40.00T* pretraining* yes [288] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 3000B* 40.00T* pretraining* yes [292] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 117B [297] 2.34T* pretraining* no [298] standard_gpu_cluster* 1.0000 0.9000 1.0500
openai 21B [303] 0.42T* pretraining* no [304] standard_gpu_cluster* 1.0000 0.9000 1.0500
openai 320B* 6.40T* pretraining* yes [308] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
openai 200B* 1.90T* pretraining* no [312] modern_hyperscale_gpu* 1.0000 1.0000 0.9000
openai 320B* 6.00T* pretraining* no [316] modern_hyperscale_gpu* 1.0000 1.0000 0.9000
xai 314B [321] 6.28T* pretraining [322] yes [323] modern_hyperscale_gpu 1.0000 1.0350 0.9000
xai 78.5B* 1.80T* pretraining* no [327] modern_hyperscale_gpu* 1.0000 1.0000 0.9000
xai 270B [329] 5.40T* pretraining* yes [323] modern_hyperscale_gpu 1.0000 1.0350 0.9000
xai 600B [41] 12.00T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
xai 96.75B* 2.40T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
xai 96.75B* 2.40T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
xai 600B* 13.00T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
xai 600B* 13.00T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000
xai 4000B* 13.00T* pretraining* yes [327] modern_hyperscale_gpu* 1.0000 1.1500 0.9000

* indicates a value retained from the project's screening rather than directly linked to an external source in this table.

Numbered source list for retained screening characteristics

The numbered references used in the retained inference and training screening-characteristic tables are listed below.

  1. [1] Active parameters. AI21 Jurassic-1 announcement
  2. [2] Context window. AI21 admissions
  3. [3] Vision. AI21 blog
  4. [4] Serving mode. AI21 admissions
  5. [5] Training parameters. AI21 Jurassic-1 announcement
  6. [6] Training modality. Jurassic-1 platform blog describing text API
  7. [7] Active parameters. Qwen2.5 32B model card
  8. [8] Context window. Qwen2.5 32B model card
  9. [9] Vision. Qwen2.5 32B model card
  10. [10] Serving mode. Qwen2.5 32B model card
  11. [11] Training parameters. Qwen2.5 32B model card
  12. [12] Training modality. Qwen2.5 32B model card
  13. [13] Active parameters. Qwen2.5 72B model card
  14. [14] Context window. Qwen2.5 72B model card
  15. [15] Vision. Qwen2.5 72B model card
  16. [16] Serving mode. Qwen2.5 72B model card
  17. [17] Training parameters. Qwen2.5 72B model card
  18. [18] Training modality. Qwen2.5 72B model card
  19. [19] Active parameters. Qwen2.5 7B model card
  20. [20] Context window. Qwen2.5 7B model card
  21. [21] Vision. Qwen2.5 7B model card
  22. [22] Serving mode. Qwen2.5 7B model card
  23. [23] Training parameters. Qwen2.5 7B model card
  24. [24] Training modality. Qwen2.5 7B model card
  25. [25] Active parameters. Qwen3 235B A22B release repository
  26. [26] Context window. Qwen3 235B A22B release repository
  27. [27] Vision. Qwen3 235B A22B release repository
  28. [28] Serving mode. Qwen3 235B A22B release repository
  29. [29] Training parameters. Qwen3 235B A22B release repository
  30. [30] Training modality. Qwen3 235B A22B release repository
  31. [31] Active parameters. Qwen3 32B release repository
  32. [32] Context window. Qwen3 32B release repository
  33. [33] Vision. Qwen3 32B release repository
  34. [34] Serving mode. Qwen3 32B release repository
  35. [35] Training parameters. Qwen3 32B release repository
  36. [36] Training modality. Qwen3 32B release repository
  37. [37] Active parameters. Third-party screening estimate from Alan D. Thompson Models Table
  38. [38] Context window. Anthropic Claude overview
  39. [39] Vision. Anthropic Claude overview
  40. [40] Serving mode. Anthropic Claude overview
  41. [41] Training parameters. Third-party screening estimate from Alan D. Thompson Models Table
  42. [42] Training modality. https://docs.anthropic.com/en/docs/about-claude/models/overview
  43. [43] Active parameters. Partial-data donor prior from source-linked market models
  44. [44] Context window. Models overview - Claude API Docs
  45. [45] Vision. Models overview - Claude API Docs
  46. [46] Serving mode. Models overview - Claude API Docs
  47. [47] Training modality. Models overview - Claude API Docs
  48. [48] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
  49. [49] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
  50. [50] Active parameters. Third-party screening estimate from Alan D. Thompson Models Table
  51. [51] Training parameters. Third-party screening estimate from Alan D. Thompson Models Table
  52. [52] Active parameters. DeepMind Gopher paper
  53. [53] Context window. DeepMind blog
  54. [54] Vision. DeepMind blog
  55. [55] Serving mode. DeepMind blog
  56. [56] Training parameters. DeepMind Gopher paper
  57. [57] Training modality. DeepMind Gopher paper centered on text modeling
  58. [58] Active parameters. DeepSeek-R1 model card
  59. [59] Context window. deepseek-ai/DeepSeek-R1
  60. [60] Vision. deepseek-ai/DeepSeek-R1
  61. [61] Serving mode. deepseek-ai/DeepSeek-R1
  62. [62] Training parameters. DeepSeek-R1 model card
  63. [63] Training modality. deepseek-ai/DeepSeek-R1
  64. [64] Active parameters. DeepSeek-V3 model card
  65. [65] Context window. deepseek-ai/DeepSeek-V3
  66. [66] Vision. deepseek-ai/DeepSeek-V3
  67. [67] Serving mode. deepseek-ai/DeepSeek-V3
  68. [68] Training parameters. DeepSeek-V3 model card
  69. [69] Training modality. deepseek-ai/DeepSeek-V3
  70. [70] Active parameters. DeepSeek V4 Flash official release note
  71. [71] Context window. DeepSeek V4 Flash documentation | DeepSeek
  72. [72] Vision. DeepSeek V4 Flash documentation | DeepSeek
  73. [73] Serving mode. DeepSeek V4 Flash documentation | DeepSeek
  74. [74] Training parameters. DeepSeek V4 Flash official release note
  75. [75] Training modality. DeepSeek V4 Flash documentation | DeepSeek
  76. [76] Active parameters. DeepSeek V4 Pro official release note
  77. [77] Context window. DeepSeek V4 Pro documentation | DeepSeek
  78. [78] Vision. DeepSeek V4 Pro documentation | DeepSeek
  79. [79] Serving mode. DeepSeek V4 Pro documentation | DeepSeek
  80. [80] Training parameters. DeepSeek V4 Pro official release note
  81. [81] Training modality. DeepSeek V4 Pro documentation | DeepSeek
  82. [82] Context window. Gemini 1.5 Flash | Gemini API
  83. [83] Vision. Gemini 1.5 Flash | Gemini API
  84. [84] Serving mode. Gemini 1.5 Flash | Gemini API
  85. [85] Training modality. Gemini 1.5 Flash | Gemini API
  86. [86] Context window. Gemini 1.5 Pro | Gemini API
  87. [87] Vision. Gemini 1.5 Pro | Gemini API
  88. [88] Serving mode. Gemini 1.5 Pro | Gemini API
  89. [89] Training modality. Gemini 1.5 Pro | Gemini API
  90. [90] Context window. Gemini models | Gemini API
  91. [91] Vision. Gemini models | Gemini API
  92. [92] Serving mode. Gemini models | Gemini API
  93. [93] Training modality. Gemini models | Gemini API
  94. [94] Active parameters. Third-party family-level screening estimate from Alan D. Thompson Models Table
  95. [95] Training parameters. Third-party family-level screening estimate from Alan D. Thompson Models Table
  96. [96] Context window. Gemini 3 Flash Preview | Gemini API
  97. [97] Vision. Gemini 3 Flash Preview | Gemini API
  98. [98] Serving mode. Gemini 3 Flash Preview | Gemini API
  99. [99] Training modality. Gemini 3 Flash Preview | Gemini API
  100. [100] Context window. Gemini 3 Pro Preview | Gemini API
  101. [101] Vision. Gemini 3 Pro Preview | Gemini API
  102. [102] Serving mode. Gemini 3 Pro Preview | Gemini API
  103. [103] Training modality. Gemini 3 Pro Preview | Gemini API
  104. [104] Context window. Gemini 3.1 Flash Live Preview | Gemini API
  105. [105] Vision. Gemini 3.1 Flash Live Preview | Gemini API
  106. [106] Serving mode. Gemini 3.1 Flash Live Preview | Gemini API
  107. [107] Training modality. Gemini 3.1 Flash Live Preview | Gemini API
  108. [108] Context window. Gemini 3.1 Flash TTS Preview | Gemini API
  109. [109] Vision. Gemini 3.1 Flash TTS Preview | Gemini API
  110. [110] Serving mode. Gemini 3.1 Flash TTS Preview | Gemini API
  111. [111] Training modality. Gemini 3.1 Flash TTS Preview | Gemini API
  112. [112] Context window. Gemini 3.1 Flash-Lite Preview | Gemini API
  113. [113] Vision. Gemini 3.1 Flash-Lite Preview | Gemini API
  114. [114] Serving mode. Gemini 3.1 Flash-Lite Preview | Gemini API
  115. [115] Training modality. Gemini 3.1 Flash-Lite Preview | Gemini API
  116. [116] Context window. Gemini 3.1 Pro Preview | Gemini API
  117. [117] Vision. Gemini 3.1 Pro Preview | Gemini API
  118. [118] Serving mode. Gemini 3.1 Pro Preview | Gemini API
  119. [119] Training modality. Gemini 3.1 Pro Preview | Gemini API
  120. [120] Active parameters. Google AI blog
  121. [121] Context window. Google AI blog
  122. [122] Vision. Google AI blog
  123. [123] Serving mode. Google AI blog
  124. [124] Training parameters. Google AI blog
  125. [125] Training modality. GLaM blog details conditional computation for text
  126. [126] Active parameters. LaMDA research paper / Google Research summary
  127. [127] Context window. Google blog
  128. [128] Vision. Google blog
  129. [129] Serving mode. Google blog
  130. [130] Training parameters. LaMDA research paper / Google Research summary
  131. [131] Training modality. LaMDA announcement describing dialog-only model
  132. [132] Active parameters. Meta Llama 3.1 model card
  133. [133] Context window. meta-llama/Llama-3.1 model card
  134. [134] Vision. meta-llama/Llama-3.1 model card
  135. [135] Serving mode. meta-llama/Llama-3.1 model card
  136. [136] Training parameters. Meta Llama 3.1 model card
  137. [137] Training modality. meta-llama/Llama-3.1 model card
  138. [138] Context window. meta-llama/Llama-3.1 model card
  139. [139] Vision. meta-llama/Llama-3.1 model card
  140. [140] Serving mode. meta-llama/Llama-3.1 model card
  141. [141] Training modality. meta-llama/Llama-3.1 model card
  142. [142] Context window. meta-llama/Llama-3.1 model card
  143. [143] Vision. meta-llama/Llama-3.1 model card
  144. [144] Serving mode. meta-llama/Llama-3.1 model card
  145. [145] Training modality. meta-llama/Llama-3.1 model card
  146. [146] Active parameters. Llama 4 Maverick model card
  147. [147] Context window. Llama 4 Maverick model card
  148. [148] Vision. Llama 4 Maverick model card
  149. [149] Serving mode. Llama 4 Maverick model card
  150. [150] Training parameters. Llama 4 Maverick model card
  151. [151] Training tokens. Official model card training-token disclosure (22T tokens)
  152. [152] Training modality. Llama 4 Maverick model card
  153. [153] Active parameters. Llama 4 Scout model card
  154. [154] Context window. Llama 4 Scout model card
  155. [155] Vision. Llama 4 Scout model card
  156. [156] Serving mode. Llama 4 Scout model card
  157. [157] Training parameters. Llama 4 Scout model card
  158. [158] Training tokens. Official model card training-token disclosure (40T tokens)
  159. [159] Training modality. Llama 4 Scout model card
  160. [160] Active parameters. OPT release
  161. [161] Context window. Meta AI blog
  162. [162] Vision. Meta AI blog
  163. [163] Serving mode. Meta AI blog
  164. [164] Training parameters. OPT release
  165. [165] Training modality. OPT release blog describing dense decoder transformer
  166. [166] Active parameters. NVidia & Microsoft announcement
  167. [167] Context window. Microsoft product page
  168. [168] Vision. Microsoft announcement
  169. [169] Serving mode. Microsoft product page
  170. [170] Training parameters. NVidia & Microsoft announcement
  171. [171] Training modality. Azure AI service introduction describes a text-only transformer
  172. [172] Active parameters. Mistral model overview
  173. [173] Context window. Codestral - Mistral Docs
  174. [174] Vision. Codestral - Mistral Docs
  175. [175] Serving mode. Codestral - Mistral Docs
  176. [176] Training parameters. Mistral model overview
  177. [177] Training modality. Codestral - Mistral Docs
  178. [178] Active parameters. Devstral 2 model documentation
  179. [179] Context window. Devstral 2 model documentation | Mistral
  180. [180] Vision. Devstral 2 model documentation | Mistral
  181. [181] Serving mode. Devstral 2 model documentation | Mistral
  182. [182] Training parameters. Devstral 2 model documentation
  183. [183] Training modality. Devstral 2 model documentation | Mistral
  184. [184] Active parameters. Ministral 3 14B model documentation
  185. [185] Context window. Ministral 3 14B model documentation | Mistral
  186. [186] Vision. Ministral 3 14B model documentation | Mistral
  187. [187] Serving mode. Ministral 3 14B model documentation | Mistral
  188. [188] Training parameters. Ministral 3 14B model documentation
  189. [189] Training modality. Ministral 3 14B model documentation | Mistral
  190. [190] Context window. Research-licensed edge model - Mistral Docs
  191. [191] Vision. Research-licensed edge model - Mistral Docs
  192. [192] Serving mode. Research-licensed edge model - Mistral Docs
  193. [193] Training modality. Research-licensed edge model - Mistral Docs
  194. [194] Context window. Research-licensed edge model - Mistral Docs
  195. [195] Vision. Research-licensed edge model - Mistral Docs
  196. [196] Serving mode. Research-licensed edge model - Mistral Docs
  197. [197] Training modality. Research-licensed edge model - Mistral Docs
  198. [198] Active parameters. Mistral Large 2 launch note
  199. [199] Context window. Mistral Large 2.1 - Mistral Docs
  200. [200] Vision. Mistral Large 2.1 - Mistral Docs
  201. [201] Serving mode. Mistral Large 2.1 - Mistral Docs
  202. [202] Training parameters. Mistral Large 2 launch note
  203. [203] Training modality. Mistral Large 2.1 - Mistral Docs
  204. [204] Active parameters. Mistral Large 3 model documentation
  205. [205] Context window. Mistral Large 3 model documentation | Mistral
  206. [206] Vision. Mistral Large 3 model documentation | Mistral
  207. [207] Serving mode. Mistral Large 3 model documentation | Mistral
  208. [208] Training parameters. Mistral Large 3 model documentation
  209. [209] Training modality. Mistral Large 3 model documentation | Mistral
  210. [210] Active parameters. Third-party screening estimate from Artificial Analysis
  211. [211] Context window. Mistral Medium 3.5 model documentation | Mistral
  212. [212] Vision. Mistral Medium 3.5 model documentation | Mistral
  213. [213] Serving mode. Mistral Medium 3.5 model documentation | Mistral
  214. [214] Training parameters. Third-party screening estimate from Artificial Analysis
  215. [215] Training modality. Mistral Medium 3.5 model documentation | Mistral
  216. [216] Context window. Mistral Small 3.1 - Mistral Docs
  217. [217] Vision. Mistral Small 3.1 - Mistral Docs
  218. [218] Serving mode. Mistral Small 3.1 - Mistral Docs
  219. [219] Training modality. Mistral Small 3.1 - Mistral Docs
  220. [220] Active parameters. Mistral Small 4 model documentation
  221. [221] Context window. Mistral Small 4 model documentation | Mistral
  222. [222] Vision. Mistral Small 4 model documentation | Mistral
  223. [223] Serving mode. Mistral Small 4 model documentation | Mistral
  224. [224] Training parameters. Mistral Small 4 model documentation
  225. [225] Training modality. Mistral Small 4 model documentation | Mistral
  226. [226] Active parameters. Third-party screening estimate from Nexos
  227. [227] Context window. OpenAI ChatGPT blog
  228. [228] Vision. OpenAI ChatGPT blog
  229. [229] Serving mode. OpenAI ChatGPT blog
  230. [230] Training parameters. Third-party screening estimate from Nexos
  231. [231] Training modality. OpenAI ChatGPT blog
  232. [232] Training hardware. OpenAI ChatGPT blog
  233. [233] Active parameters. Third-party parameter estimate from Exploding Topics
  234. [234] Context window. OpenAI GPT-4 research release
  235. [235] Vision. OpenAI GPT-4 research release
  236. [236] Serving mode. OpenAI GPT-4 research release
  237. [237] Training parameters. Third-party parameter estimate from Exploding Topics
  238. [238] Training regime. OpenAI GPT-4 technical report
  239. [239] Training modality. OpenAI GPT-4 research release
  240. [240] Training hardware. OpenAI GPT-4 research release
  241. [241] Context window. GPT-4o model docs | OpenAI API
  242. [242] Vision. GPT-4o model docs | OpenAI API
  243. [243] Serving mode. GPT-4o model docs | OpenAI API
  244. [244] Training modality. GPT-4o model docs | OpenAI API
  245. [245] Context window. GPT-4o mini model docs | OpenAI API
  246. [246] Vision. Introducing GPT-4o mini | OpenAI
  247. [247] Serving mode. GPT-4o mini model docs | OpenAI API
  248. [248] Training modality. Introducing GPT-4o mini | OpenAI
  249. [249] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
  250. [250] Context window. GPT-5 mini Model | OpenAI API
  251. [251] Vision. GPT-5 mini Model | OpenAI API
  252. [252] Serving mode. GPT-5 mini Model | OpenAI API
  253. [253] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
  254. [254] Training modality. GPT-5 mini Model | OpenAI API
  255. [255] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
  256. [256] Context window. GPT-5 nano Model | OpenAI API
  257. [257] Vision. GPT-5 nano Model | OpenAI API
  258. [258] Serving mode. GPT-5 nano Model | OpenAI API
  259. [259] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
  260. [260] Training modality. GPT-5 nano Model | OpenAI API
  261. [261] Context window. GPT-5.2 model docs | OpenAI API
  262. [262] Vision. GPT-5.2 model docs | OpenAI API
  263. [263] Serving mode. GPT-5.2 model docs | OpenAI API
  264. [264] Training modality. GPT-5.2 model docs | OpenAI API
  265. [265] Context window. GPT-5.2-pro model docs | OpenAI API
  266. [266] Vision. GPT-5.2-pro model docs | OpenAI API
  267. [267] Serving mode. GPT-5.2-pro model docs | OpenAI API
  268. [268] Training modality. GPT-5.2-pro model docs | OpenAI API
  269. [269] Context window. GPT-5.4 model docs | OpenAI API
  270. [270] Vision. GPT-5.4 model docs | OpenAI API
  271. [271] Serving mode. GPT-5.4 model docs | OpenAI API
  272. [272] Training modality. GPT-5.4 model docs | OpenAI API
  273. [273] Context window. GPT-5.4 mini model docs | OpenAI API
  274. [274] Vision. GPT-5.4 mini model docs | OpenAI API
  275. [275] Serving mode. GPT-5.4 mini model docs | OpenAI API
  276. [276] Training modality. GPT-5.4 mini model docs | OpenAI API
  277. [277] Context window. GPT-5.4 nano model docs | OpenAI API
  278. [278] Vision. GPT-5.4 nano model docs | OpenAI API
  279. [279] Serving mode. GPT-5.4 nano model docs | OpenAI API
  280. [280] Training modality. GPT-5.4 nano model docs | OpenAI API
  281. [281] Context window. GPT-5.4-pro model docs | OpenAI API
  282. [282] Vision. GPT-5.4-pro model docs | OpenAI API
  283. [283] Serving mode. GPT-5.4-pro model docs | OpenAI API
  284. [284] Training modality. GPT-5.4-pro model docs | OpenAI API
  285. [285] Context window. GPT-5.5 model docs | OpenAI API
  286. [286] Vision. GPT-5.5 model docs | OpenAI API
  287. [287] Serving mode. GPT-5.5 model docs | OpenAI API
  288. [288] Training modality. GPT-5.5 model docs | OpenAI API
  289. [289] Context window. GPT-5.5-pro model docs | OpenAI API
  290. [290] Vision. GPT-5.5-pro model docs | OpenAI API
  291. [291] Serving mode. GPT-5.5-pro model docs | OpenAI API
  292. [292] Training modality. GPT-5.5-pro model docs | OpenAI API
  293. [293] Active parameters. gpt-oss-120b Model | OpenAI API
  294. [294] Context window. gpt-oss-120b Model | OpenAI API
  295. [295] Vision. gpt-oss-120b Model | OpenAI API
  296. [296] Serving mode. gpt-oss-120b Model | OpenAI API
  297. [297] Training parameters. gpt-oss-120b Model | OpenAI API
  298. [298] Training modality. gpt-oss-120b Model | OpenAI API
  299. [299] Active parameters. gpt-oss-20b Model | OpenAI API
  300. [300] Context window. gpt-oss-20b Model | OpenAI API
  301. [301] Vision. gpt-oss-20b Model | OpenAI API
  302. [302] Serving mode. gpt-oss-20b Model | OpenAI API
  303. [303] Training parameters. gpt-oss-20b Model | OpenAI API
  304. [304] Training modality. gpt-oss-20b Model | OpenAI API
  305. [305] Context window. o1 model docs | OpenAI API
  306. [306] Vision. o1 model docs | OpenAI API
  307. [307] Serving mode. o1 model docs | OpenAI API
  308. [308] Training modality. o1 model docs | OpenAI API
  309. [309] Context window. o1-mini model docs | OpenAI API
  310. [310] Vision. o1-mini model docs | OpenAI API
  311. [311] Serving mode. o1-mini model docs | OpenAI API
  312. [312] Training modality. o1-mini model docs | OpenAI API
  313. [313] Context window. o1-preview model docs | OpenAI API
  314. [314] Vision. o1-preview model docs | OpenAI API
  315. [315] Serving mode. o1-preview model docs | OpenAI API
  316. [316] Training modality. o1-preview model docs | OpenAI API
  317. [317] Active parameters. Open Release of Grok-1 | xAI
  318. [318] Context window. https://docs.x.ai/developers/models
  319. [319] Vision. https://docs.x.ai/developers/models
  320. [320] Serving mode. https://docs.x.ai/developers/models
  321. [321] Training parameters. Open Release of Grok-1 | xAI
  322. [322] Training regime. Open Release of Grok-1 | xAI
  323. [323] Training modality. https://docs.x.ai/developers/models
  324. [324] Context window. Models and Pricing | xAI
  325. [325] Vision. Models and Pricing | xAI
  326. [326] Serving mode. Models and Pricing | xAI
  327. [327] Training modality. Models and Pricing | xAI
  328. [328] Active parameters. Third-party screening estimate from the Grok-2 open-weight config discussion
  329. [329] Training parameters. Third-party screening estimate from the Grok-2 open-weight config discussion

Country factors for carbon and water recalculation

No. Country Year Carbon intensity Reference
1 France 2024 40 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.
2 United States 2024 385 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.
3 Germany 2024 380 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.
4 United Kingdom 2024 180 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.
5 Canada 2024 120 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.
6 China 2024 540 gCO2e/kWh ImpactLLM method note: project screening electricity factors
Project screening default retained for rapid comparative estimation; not an audited national inventory factor.