An Open Tool for Estimating the Environmental Footprint of LLMs
Comparative environmental impact of models
This public benchmark currently quantifies 83 market models. 56 use a strict parameter basis (directly sourced, directly derivable, or sourced from an explicit third-party estimate), while 27 additional recent models use a documented partial-data donor prior derived from the strict catalog.
The chart below shows the estimated central values for the 83 quantified catalog models under a standardized inference scenario corresponding to 1 hour of active use: 34.6 interactions/hour, 1000 input tokens, 550 output tokens, and one LLM request per use. The hourly pace is derived from an average reading speed of 238.0 words/min (Brysbaert, 2019) and a project convention of 1 token ≈ 0.75 word.
Benchmarks integrated into the chart now include direct household-use examples from Purdue Extension (fluorescent lamp ≈ 9.3 Wh over 1 h; laptop ≈ 32 Wh over 1 h), plus explicit project-scale equivalents for everyday uses better aligned with the current LLM range: laptop use for 3 h ≈ 96 Wh, laptop use for 5 h ≈ 160 Wh, electric kettle for 7 min ≈ 210 Wh, and electric space heater for 10 min ≈ 250 Wh. For carbon, the chart uses an average gasoline car benchmark derived from the ICCT (2025) factor retained by the project (235 gCO2e/km), shown here both for 430 m ≈ 101.1 gCO2e and for 1 km ≈ 235 gCO2e. Exact retained benchmark values are listed in the annex.
Inference vs. training impact map
This scatter plot compares each quantified model on two axes at once: standardized inference impact over one hour on the horizontal axis and retained training impact on the vertical axis. Point size follows the retained active parameter basis, while colors distinguish providers.
Inference bubble chart
This explicit bubble chart positions each model by effective active parameters and by its retained inference impact. Bubble size reflects the retained context window, while colors distinguish providers.
Inference uncertainty span by model
This chart makes the project’s retained inference range explicit for each model by showing the low, central, and high values under the standardized one-hour scenario.
Inference model landscape
This landscape view clusters the quantified models from the characteristics retained by the project for inference screening: active and effective parameter basis, context window, serving mode, modality support, architecture notes, and central energy and carbon outputs. Nearby points indicate models with similar retained screening profiles, not a simple one-metric ranking.
Inference screening factor heatmap
This heatmap exposes the central screening factors retained for each quantified market model. It shows the four multiplicative factors used by the project’s prompt proxy and the resulting ratio between effective and raw active parameters.
Inference carbon vs. parameter count
This complementary view places models by retained active parameter count on the horizontal axis and by central inference carbon over one hour on the vertical axis, using logarithmic scaling on both axes.
Country-mix sensitivity
This view compares central inference energy and carbon over one hour, while coloring each model by the retained electricity-mix country used for carbon recalculation. It helps separate model-size effects from country-mix effects.
Inference carbon by model release date
This timeline follows the evolution of the project’s central inference CO2e estimate over time for the OpenAI, Claude, Grok, and Mistral families, using the release month of each model as the horizontal axis.
Inference CO2 doubling view
This discussion-oriented chart summarizes the central inference screening values for flagship GPT, Claude, and Grok models as a simple doubling-time reading. It should be read as an interpretation of the retained observatory values, not as a provider-side measurement law.
Under the current central screening profile, the flagship inference series suggests a slower increase than training, because standardized usage, active compute, and per-request serving assumptions damp part of the growth that appears in total model scale.
Comparative training impacts of models
This public training benchmark currently quantifies 83 market models: 56 with a strict retained parameter basis and 27 with a documented partial-data donor prior.
The chart below shows the central values retained for the quantified models across two training indicator families: training energy and training CO2e contextualized from retained training energy and the model country proxy. The current screening method combines retained parameter count, a training-token prior, a training-regime prior, architecture features, and a hardware-class proxy. Everyday benchmarks are inserted directly into the list to situate those scales, not to imply direct observed equivalence.
Benchmarks integrated into the chart: household electricity for 74,277 households over one year of domestic use, i.e. ≈ 0.19 TWh based on an average consumption of 2,500 kWh per household (RTE, 2021 estimate), and full-flight aviation derived from Klöwer et al. (2025) from 577.97 MtCO2 and 27.45 million commercial flights observed in 2023, i.e. ≈ 71,491.6 tCO2e for 3,396 full flights. These comparison points are aligned with the current maximum central screening order of magnitude in the training chart, not with a direct provider-side measurement.
Training model landscape
This landscape view clusters the quantified models from the characteristics retained by the project for training screening: retained parameter basis, training-token prior, training regime, hardware-class proxy, modality support, architecture notes, and central training energy and carbon outputs. Nearby points indicate similar retained screening profiles rather than a direct ranking on one axis.
Training screening factor heatmap
This heatmap exposes the central screening factors retained for each quantified market model in the training proxy. It shows the regime, architecture, and hardware factors together with the retained training-token ratio per parameter.
Training uncertainty span by model
This view shows the low, central, and high training CO2e values contextualized from retained training energy for each quantified market model. It makes explicit how widely the training proxy can vary once the parameter and token exponents, donor priors, and contextual factors are widened.
Training carbon vs. parameter count
This complementary view places models by retained parameter count on the horizontal axis and by contextualized training CO2e on the vertical axis, using logarithmic scaling on both axes.
Training CO2e by model release date
This timeline follows the evolution of the project’s retained training CO2e estimate over time for the OpenAI, Claude, Grok, and Mistral families, using the release month of each model as the horizontal axis.
Training CO2e doubling view
This chart compresses the central training screening values of flagship GPT, Claude, and Grok models into a simple doubling-time interpretation. It is meant as a discussion support to make structural acceleration legible, not as a claim of direct industrial telemetry.
The apparent acceleration is stronger for training because the current screening method compounds retained parameter count, token priors, architecture effects, and hardware assumptions. The resulting doubling pace is therefore a transparent scenario reading, not a universal empirical constant.
83 market models quantified by the project
The table below compares the market models retained in the public quantitative benchmark under the same inference scenario. For each model, the application shows the central values produced by the project’s multi-factor prompt proxy, both per hour of standardized use and per request. The current catalog combines 56 strict source-linked rows and 27 partial-data rows derived from donor models with comparable sourced metadata.
| 100B* | US Screening proxy |
6.0 Wh | 2.3 gCO2e | 0.17 Wh | 0.0669 gCO2e | |
| 100B* | US Screening proxy |
6.0 Wh | 2.3 gCO2e | 0.17 Wh | 0.0669 gCO2e | |
| 100B* | US Screening proxy |
6.0 Wh | 2.3 gCO2e | 0.17 Wh | 0.0669 gCO2e | |
| 100B* | US Screening proxy |
6.0 Wh | 2.3 gCO2e | 0.17 Wh | 0.0669 gCO2e | |
| 22B* | US Screening proxy |
1.4 Wh | 0.55 gCO2e | 0.0412 Wh | 0.0159 gCO2e | |
| 175B* | US Screening proxy |
10.2 Wh | 3.9 gCO2e | 0.30 Wh | 0.11 gCO2e | |
| 175B* | US Screening proxy |
10.2 Wh | 3.9 gCO2e | 0.30 Wh | 0.11 gCO2e | |
| 5000B* | US Screening proxy |
264.7 Wh | 101.9 gCO2e | 7.7 Wh | 2.9 gCO2e | |
| 2000B* | US Screening proxy |
103.5 Wh | 39.9 gCO2e | 3.0 Wh | 1.2 gCO2e | |
| 5000B* | US Screening proxy |
264.7 Wh | 101.9 gCO2e | 7.7 Wh | 2.9 gCO2e | |
| 4000B* | US Screening proxy |
214.1 Wh | 82.4 gCO2e | 6.2 Wh | 2.4 gCO2e | |
| 175B* | US Screening proxy |
10.2 Wh | 3.9 gCO2e | 0.30 Wh | 0.11 gCO2e | |
| 400B* | US Screening proxy |
24.0 Wh | 9.2 gCO2e | 0.69 Wh | 0.27 gCO2e | |
| 22B | FR Provider-country proxy |
1.2 Wh | 0.0484 gCO2e | 0.0347 Wh | 0.0014 gCO2e | |
| 37B active / 671B total | CN Provider-country proxy |
2.9 Wh | 1.6 gCO2e | 0.0836 Wh | 0.0451 gCO2e | |
| 37B active / 671B total | CN Provider-country proxy |
2.8 Wh | 1.5 gCO2e | 0.0798 Wh | 0.0431 gCO2e | |
| 13B active / 284B total | CN Provider-country proxy |
1.1 Wh | 0.61 gCO2e | 0.0327 Wh | 0.0177 gCO2e | |
| 49B active / 1600B total | CN Provider-country proxy |
4.3 Wh | 2.3 gCO2e | 0.12 Wh | 0.0675 gCO2e | |
| 123B | FR Provider-country proxy |
6.8 Wh | 0.27 gCO2e | 0.20 Wh | 0.0078 gCO2e | |
| 133.5B* | US Screening proxy |
8.5 Wh | 3.3 gCO2e | 0.25 Wh | 0.0944 gCO2e | |
| 133.5B* | US Screening proxy |
8.5 Wh | 3.3 gCO2e | 0.25 Wh | 0.0944 gCO2e | |
| 30B* | US Screening proxy |
2.1 Wh | 0.79 gCO2e | 0.0594 Wh | 0.0229 gCO2e | |
| 80B* | US Screening proxy |
5.2 Wh | 2.0 gCO2e | 0.15 Wh | 0.0581 gCO2e | |
| 30B* | US Screening proxy |
2.1 Wh | 0.79 gCO2e | 0.0594 Wh | 0.0229 gCO2e | |
| 200B* | US Screening proxy |
12.5 Wh | 4.8 gCO2e | 0.36 Wh | 0.14 gCO2e | |
| 30B* | US Screening proxy |
2.1 Wh | 0.79 gCO2e | 0.0594 Wh | 0.0229 gCO2e | |
| 3000B* | US Screening proxy |
163.2 Wh | 62.8 gCO2e | 4.7 Wh | 1.8 gCO2e | |
| 30B* | US Screening proxy |
1.9 Wh | 0.72 gCO2e | 0.0543 Wh | 0.0209 gCO2e | |
| 22B* | US Screening proxy |
1.3 Wh | 0.49 gCO2e | 0.0369 Wh | 0.0142 gCO2e | |
| 22B* | US Screening proxy |
1.5 Wh | 0.59 gCO2e | 0.0442 Wh | 0.0170 gCO2e | |
| 3000B* | US Screening proxy |
163.2 Wh | 62.8 gCO2e | 4.7 Wh | 1.8 gCO2e | |
| 130B | US Documented region, country retained as reference |
6.9 Wh | 2.7 gCO2e | 0.20 Wh | 0.0768 gCO2e | |
| 280B | GB Documented region, country retained as reference |
14.3 Wh | 2.6 gCO2e | 0.41 Wh | 0.0744 gCO2e | |
| 175B* | US Screening proxy |
9.2 Wh | 3.5 gCO2e | 0.26 Wh | 0.10 gCO2e | |
| 440B* active / 1760B* total | US Screening proxy |
26.0 Wh | 10.0 gCO2e | 0.75 Wh | 0.29 gCO2e | |
| 200B* | US Screening proxy |
11.4 Wh | 4.4 gCO2e | 0.33 Wh | 0.13 gCO2e | |
| 200B* | US Screening proxy |
11.4 Wh | 4.4 gCO2e | 0.33 Wh | 0.13 gCO2e | |
| 95B* | US Screening proxy |
5.9 Wh | 2.3 gCO2e | 0.17 Wh | 0.0657 gCO2e | |
| 95B* | US Screening proxy |
5.9 Wh | 2.3 gCO2e | 0.17 Wh | 0.0657 gCO2e | |
| 3000B* | US Screening proxy |
156.8 Wh | 60.4 gCO2e | 4.5 Wh | 1.7 gCO2e | |
| 440B* | US Screening proxy |
25.3 Wh | 9.7 gCO2e | 0.73 Wh | 0.28 gCO2e | |
| 3000B* | US Screening proxy |
156.8 Wh | 60.4 gCO2e | 4.5 Wh | 1.7 gCO2e | |
| 95B* | US Screening proxy |
5.9 Wh | 2.3 gCO2e | 0.17 Wh | 0.0657 gCO2e | |
| 95B* | US Screening proxy |
5.9 Wh | 2.3 gCO2e | 0.17 Wh | 0.0657 gCO2e | |
| 3000B* | US Screening proxy |
156.8 Wh | 60.4 gCO2e | 4.5 Wh | 1.7 gCO2e | |
| 3000B* | US Screening proxy |
156.8 Wh | 60.4 gCO2e | 4.5 Wh | 1.7 gCO2e | |
| 3000B* | US Screening proxy |
163.3 Wh | 62.9 gCO2e | 4.7 Wh | 1.8 gCO2e | |
| 5.1B active / 117B total | US Comparative reference country |
0.40 Wh | 0.16 gCO2e | 0.0116 Wh | 0.0045 gCO2e | |
| 3.6B active / 21B total | US Comparative reference country |
0.26 Wh | 0.10 gCO2e | 0.00740 Wh | 0.0029 gCO2e | |
| 78.5B* active / 314B* total | US Documented region, country retained as reference |
5.5 Wh | 2.1 gCO2e | 0.16 Wh | 0.0612 gCO2e | |
| 78.5B* | US Documented region, country retained as reference |
4.8 Wh | 1.8 gCO2e | 0.14 Wh | 0.0531 gCO2e | |
| 115B* active / 270B* total | US Documented region, country retained as reference |
8.3 Wh | 3.2 gCO2e | 0.24 Wh | 0.0920 gCO2e | |
| 600B* | US Documented region, country retained as reference |
38.0 Wh | 14.6 gCO2e | 1.1 Wh | 0.42 gCO2e | |
| 96.75B* | US Documented region, country retained as reference |
6.4 Wh | 2.5 gCO2e | 0.19 Wh | 0.0714 gCO2e | |
| 96.75B* | US Documented region, country retained as reference |
6.7 Wh | 2.6 gCO2e | 0.19 Wh | 0.0748 gCO2e | |
| 600B* | US Documented region, country retained as reference |
36.3 Wh | 14.0 gCO2e | 1.0 Wh | 0.40 gCO2e | |
| 600B* | US Documented region, country retained as reference |
38.0 Wh | 14.6 gCO2e | 1.1 Wh | 0.42 gCO2e | |
| 4000B* | US Documented region, country retained as reference |
220.2 Wh | 84.8 gCO2e | 6.4 Wh | 2.5 gCO2e | |
| 178B | US Documented region, country retained as reference |
9.3 Wh | 3.6 gCO2e | 0.27 Wh | 0.10 gCO2e | |
| 137B* | US Documented region, country retained as reference |
7.3 Wh | 2.8 gCO2e | 0.21 Wh | 0.0807 gCO2e | |
| 405B | US Comparative reference country |
19.1 Wh | 7.4 gCO2e | 0.55 Wh | 0.21 gCO2e | |
| 70B | US Comparative reference country |
3.6 Wh | 1.4 gCO2e | 0.10 Wh | 0.0402 gCO2e | |
| 8B | US Comparative reference country |
0.46 Wh | 0.18 gCO2e | 0.0133 Wh | 0.0051 gCO2e | |
| 17B active / 400B total | US Comparative reference country |
1.4 Wh | 0.55 gCO2e | 0.0410 Wh | 0.0158 gCO2e | |
| 17B active / 109B total | US Comparative reference country |
1.4 Wh | 0.54 gCO2e | 0.0402 Wh | 0.0155 gCO2e | |
| 530B | US Screening proxy |
26.2 Wh | 10.1 gCO2e | 0.76 Wh | 0.29 gCO2e | |
| 14B | FR Provider-country proxy |
0.88 Wh | 0.0346 gCO2e | 0.0255 Wh | 0.0010 gCO2e | |
| 3B | FR Provider-country proxy |
0.19 Wh | 0.0069 gCO2e | 0.00560 Wh | 0.0002 gCO2e | |
| 8B | FR Provider-country proxy |
0.49 Wh | 0.0208 gCO2e | 0.0142 Wh | 0.0006 gCO2e | |
| 123B | FR Provider-country proxy |
7.2 Wh | 0.29 gCO2e | 0.21 Wh | 0.0083 gCO2e | |
| 41B active / 675B total | FR Provider-country proxy |
3.2 Wh | 0.13 gCO2e | 0.0925 Wh | 0.0037 gCO2e | |
| 128B* | FR Provider-country proxy |
7.2 Wh | 0.29 gCO2e | 0.21 Wh | 0.0084 gCO2e | |
| 24B | FR Provider-country proxy |
1.4 Wh | 0.0588 gCO2e | 0.0413 Wh | 0.0017 gCO2e | |
| 6.5B active / 119B total | FR Provider-country proxy |
0.56 Wh | 0.0208 gCO2e | 0.0162 Wh | 0.0006 gCO2e | |
| 320B* | US Screening proxy |
18.2 Wh | 7.0 gCO2e | 0.52 Wh | 0.20 gCO2e | |
| 200B* | US Screening proxy |
11.6 Wh | 4.5 gCO2e | 0.34 Wh | 0.13 gCO2e | |
| 320B* | US Screening proxy |
18.1 Wh | 7.0 gCO2e | 0.52 Wh | 0.20 gCO2e | |
| 175B | US Documented region, country retained as reference |
8.1 Wh | 3.1 gCO2e | 0.23 Wh | 0.0900 gCO2e | |
| 32B | CN Provider-country proxy |
1.7 Wh | 0.93 gCO2e | 0.0496 Wh | 0.0268 gCO2e | |
| 72B | CN Provider-country proxy |
3.7 Wh | 2.0 gCO2e | 0.11 Wh | 0.0579 gCO2e | |
| 7B | CN Provider-country proxy |
0.40 Wh | 0.22 gCO2e | 0.0117 Wh | 0.0063 gCO2e | |
| 22B active / 235B total | CN Provider-country proxy |
1.5 Wh | 0.82 gCO2e | 0.0437 Wh | 0.0236 gCO2e | |
| 32.8B | CN Provider-country proxy |
1.8 Wh | 0.95 gCO2e | 0.0508 Wh | 0.0274 gCO2e |
`Retained country` is the country actually used to recalculate CO2 via the electricity mix. When the exact country is not published, the project uses an explicit screening proxy rather than presenting a location as certain.
`*` indicates an estimated parameter count rather than a provider-published value.
The market-model comparison now relies on market_multifactor_prompt_proxy_v1: a prompt-energy screening proxy whose main prompt-level calibration anchor comes from Elsworth et al. (2025), then adjusted by active parameters, context window, serving mode, modality support, architecture overhead, and standardized token volume, and interpreted alongside other inference references.
83 market models with quantified training impacts
This table projects the training orders of magnitude of the quantified market models from the indicator families actually available in the literature: training energy derived from emissions when the source country is documented in the electricity-mix table, and training CO2e contextualized from that retained energy and the model country proxy. The current screening proxy combines retained parameter count, a training-token prior, a training-regime prior, architecture features, and a hardware-class proxy. The quantified layer combines 56 strict rows and 27 partial-data donor priors.
| 100B* | 186.2 MWh | 71.71 tCO2e | |
| 100B* | 41.0 MWh | 15.78 tCO2e | |
| 100B* | 484.2 MWh | 186.4 tCO2e | |
| 100B* | 223.5 MWh | 86.05 tCO2e | |
| 22B* | 9.0 MWh | 3.47 tCO2e | |
| 175B* | 570.4 MWh | 219.6 tCO2e | |
| 175B* | 570.4 MWh | 219.6 tCO2e | |
| 5000B* | 50.9 GWh | 19 614 tCO2e | |
| 2000B* | 74.5 GWh | 28 682 tCO2e | |
| 5000B* | 98.0 GWh | 37 719 tCO2e | |
| 4000B* | 62.7 GWh | 24 140 tCO2e | |
| 175B* | 1.3 GWh | 501.9 tCO2e | |
| 400B* | 3.0 GWh | 1 147 tCO2e | |
| 22B | 8.7 MWh | 0.35 tCO2e | |
| 37B active / 671B total | 7.3 GWh | 3 938 tCO2e | |
| 37B active / 671B total | 7.3 GWh | 3 938 tCO2e | |
| 13B active / 284B total | 1.3 GWh | 705.4 tCO2e | |
| 49B active / 1600B total | 41.5 GWh | 22 389 tCO2e | |
| 123B | 272.2 MWh | 10.89 tCO2e | |
| 133.5B* | 149.2 MWh | 57.44 tCO2e | |
| 133.5B* | 298.4 MWh | 114.9 tCO2e | |
| 30B* | 16.8 MWh | 6.45 tCO2e | |
| 80B* | 119.2 MWh | 45.89 tCO2e | |
| 30B* | 16.8 MWh | 6.45 tCO2e | |
| 200B* | 745.0 MWh | 286.8 tCO2e | |
| 30B* | 55.9 MWh | 21.51 tCO2e | |
| 3000B* | 185.7 GWh | 71 492 tCO2e | |
| 30B* | 55.9 MWh | 21.51 tCO2e | |
| 22B* | 16.4 MWh | 6.31 tCO2e | |
| 22B* | 16.4 MWh | 6.31 tCO2e | |
| 3000B* | 185.7 GWh | 71 492 tCO2e | |
| 130B | 379.0 MWh | 145.9 tCO2e | |
| 280B | 1.4 GWh | 244.9 tCO2e | |
| 175B* | 578.6 MWh | 222.8 tCO2e | |
| 440B* active / 1760B* total | 8.9 GWh | 3 407 tCO2e | |
| 200B* | 745.0 MWh | 286.8 tCO2e | |
| 200B* | 29.8 MWh | 11.47 tCO2e | |
| 95B* | 168.1 MWh | 64.71 tCO2e | |
| 95B* | 168.1 MWh | 64.71 tCO2e | |
| 3000B* | 185.7 GWh | 71 492 tCO2e | |
| 440B* | 2.5 GWh | 946.5 tCO2e | |
| 3000B* | 185.7 GWh | 71 492 tCO2e | |
| 95B* | 168.1 MWh | 64.71 tCO2e | |
| 95B* | 168.1 MWh | 64.71 tCO2e | |
| 3000B* | 111.4 GWh | 42 895 tCO2e | |
| 3000B* | 123.8 GWh | 47 661 tCO2e | |
| 3000B* | 123.8 GWh | 47 661 tCO2e | |
| 5.1B active / 117B total | 232.8 MWh | 89.62 tCO2e | |
| 3.6B active / 21B total | 7.5 MWh | 2.89 tCO2e | |
| 78.5B* active / 314B* total | 1.7 GWh | 636.3 tCO2e | |
| 78.5B* | 114.4 MWh | 44.05 tCO2e | |
| 115B* active / 270B* total | 1.2 GWh | 470.5 tCO2e | |
| 600B* | 6.7 GWh | 2 581 tCO2e | |
| 96.75B* | 216.2 MWh | 83.25 tCO2e | |
| 96.75B* | 216.2 MWh | 83.25 tCO2e | |
| 600B* | 7.3 GWh | 2 797 tCO2e | |
| 600B* | 7.3 GWh | 2 797 tCO2e | |
| 4000B* | 10.2 GWh | 3 923 tCO2e | |
| 178B | 840.8 MWh | 323.7 tCO2e | |
| 137B* | 332.8 MWh | 128.1 tCO2e | |
| 405B | 3.1 GWh | 1 193 tCO2e | |
| 70B | 92.6 MWh | 35.64 tCO2e | |
| 8B | 1.1 MWh | 0.42 tCO2e | |
| 17B active / 400B total | 8.6 GWh | 3 313 tCO2e | |
| 17B active / 109B total | 4.3 GWh | 1 641 tCO2e | |
| 530B | 8.6 GWh | 3 305 tCO2e | |
| 14B | 4.5 MWh | 0.18 tCO2e | |
| 3B | 179.4 kWh | 0.01 tCO2e | |
| 8B | 1.0 MWh | 0.04 tCO2e | |
| 123B | 281.8 MWh | 11.27 tCO2e | |
| 41B active / 675B total | 8.5 GWh | 339.4 tCO2e | |
| 128B* | 339.1 MWh | 13.56 tCO2e | |
| 24B | 11.9 MWh | 0.48 tCO2e | |
| 6.5B active / 119B total | 263.7 MWh | 10.55 tCO2e | |
| 320B* | 1.9 GWh | 734.3 tCO2e | |
| 200B* | 307.7 MWh | 118.5 tCO2e | |
| 320B* | 1.6 GWh | 598.6 tCO2e | |
| 175B | 661.3 MWh | 254.6 tCO2e | |
| 32B | 19.3 MWh | 10.45 tCO2e | |
| 72B | 97.9 MWh | 52.89 tCO2e | |
| 7B | 826.0 kWh | 0.45 tCO2e | |
| 22B active / 235B total | 939.1 MWh | 507.1 tCO2e | |
| 32.8B | 20.3 MWh | 10.98 tCO2e |
`*` indicates an estimated parameter count rather than a provider-published value.
ImpactLLM is designed as a transparent screening tool, not as a black-box score. The current release starts from source-linked inference anchors, then exposes a bounded multi-factor proxy rather than a hidden single-number score.
1. Source-linked literature anchors.
The application-level estimator starts from published inference indicators linked to an explicit source, model, geography, and system boundary. In the current market-model release, the predictive core uses Elsworth et al. (2025) as the main prompt-level calibration anchor, with a median prompt energy of 0.24 Wh/prompt for Gemini Apps, and is interpreted alongside other inference references such as the ML.ENERGY Benchmark, Ren et al. (2024), and Li et al. (2025).
2. A multi-factor effective-parameter proxy.
When direct telemetry is unavailable for a target model, ImpactLLM does not rely on a raw parameter multiple alone. It builds an effective active-parameter profile from the retained model characteristics: active parameters, context window, serving mode (open, hybrid, closed), modality support, and architecture notes such as MoE or reasoning-oriented overheads.
3. Token volume remains explicit.
The current proxy adjusts the anchor with a weighted prompt-compute volume defined from input and output tokens. Output generation is weighted more heavily than input processing, so output-heavy scenarios and repeated LLM calls raise the estimate materially.
The current prompt-level branch is a screening proxy, not an audited benchmark. For this reason, the application returns a bounded low-central-high result rather than one falsely precise deterministic value.
4. Carbon derived from context.
Carbon is not copied mechanically from the source paper. It is recalculated from the retained energy estimate using the electricity mix associated with the selected country context.
5. A research-oriented estimator.
The result is an auditable estimate intended for comparison, software design, and methodological discussion. It is useful precisely because the assumptions, factors, and retained sources remain visible and inspectable.
Technical paper
Pachot, A., & Petit, T. (2026, March 14). Transparent Screening for LLM Inference and Training Impacts. /impact-llm/downloads/ImpactLLM_paper.pdf
We work on responsible AI with a focus on methodological rigor, traceability, and real-world decision support. Our work combines scientific research, product design, and operational deployment to make AI systems more transparent, more accountable, and more useful in practice.
How to cite ImpactLLM
Pachot, A., & Petit, T. (2026, March 14). Transparent Screening for LLM Inference and Training Impacts.
BibTeX
@misc{impactllm_screening_2026,
title = {Transparent Screening for LLM Inference and Training Impacts},
author = {Pachot, Arnault and Petit, Thierry},
year = {2026},
month = mar,
note = {Conference paper preprint},
url = {https://dev.emotia.com/impact-llm/downloads/ImpactLLM_paper.pdf}
}
Download technical paper PDF | Download technical paper BibTeX
GitHub repository
The project repository is available on GitHub: https://github.com/apachot/ImpactLLM.
License
This program is free software: you can redistribute it and/or modify it under the terms of the GNU General Public License as published by the Free Software Foundation, either version 3 of the License, or (at your option) any later version.
This program is distributed in the hope that it will be useful, but WITHOUT ANY WARRANTY; without even the implied warranty of MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU General Public License for more details.
You should have received a copy of the GNU General Public License along with this program. If not, see https://www.gnu.org/licenses/.
Arnault Pachot
Arnault Pachot is a researcher and entrepreneur, founder of OpenStudio and now founder of Emotia. He works on responsible digital transformation, Green IT, and decision-oriented AI systems. He co-authored the Dunod book Intelligence artificielle et environnement : alliance ou nuisance ?, dedicated to practical pathways for environmentally responsible AI.
Google Scholar: Arnault Pachot
Thierry Petit
Thierry Petit is a senior AI researcher and scientific leader with more than twenty years of academic and R&D experience in Europe and the United States. His work spans trustworthy AI, simulation, optimization, and decision-grade platforms. At Emotia and Pollitics, he leads the scientific direction of systems designed to remain both operationally useful and methodologically robust.
Selected references on AI and the environment
- Pachot, A., Patissier, C., & Open Studio. (2022). Intelligence artificielle et environnement : alliance ou nuisance ? L'IA face aux défis écologiques d'aujourd'hui et de demain. Dunod. https://www.dunod.com/entreprise-et-economie/intelligence-artificielle-et-environnement-alliance-ou-nuisance-ia-face-aux
- Pachot, A., & Patissier, C. (2023). Toward Sustainable Artificial Intelligence: An Overview of Environmental Protection Uses and Issues. Green and Low-Carbon Economy, 3(2), 105-112. https://ojs.bonviewpress.com/index.php/GLCE/article/view/608
This annex brings together the quantified reference material used in the interface, along with everyday comparison benchmarks and country factors used for carbon and water recalculation.
Inference reference set
Training reference set
| Ref. | Data type | LLM model | Parameters | Country | Value | Citation |
|---|---|---|---|---|---|---|
| [1] | Greenhouse gas emissions from training (Transformer (big)) | Transformer (big) | 213M | États-Unis | 192 lb CO2e | Strubell, E., Ganesh, A., & McCallum, A. (2019). Energy and Policy Considerations for Deep Learning in NLP. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 3645--3650. Table 1, p. 1 |
| [2] | Greenhouse gas emissions from training (BLOOM 176B) | BLOOM 176B | 176B | France | 24.7 tCO2e | Luccioni, A. S., Viguier, S., & Ligozat, A. L. (2023). Estimating the Carbon Footprint of BLOOM. Journal of Machine Learning Research, 24(253), 1--15. https://www.jmlr.org/papers/v24/23-0069.html Abstract, p. 1 |
| [3] | Greenhouse gas emissions from training (BLOOM 176B) | BLOOM 176B | 176B | France | 50.5 tCO2e | Luccioni, A. S., Viguier, S., & Ligozat, A. L. (2023). Estimating the Carbon Footprint of BLOOM. Journal of Machine Learning Research, 24(253), 1--15. https://www.jmlr.org/papers/v24/23-0069.html Abstract, p. 1; Table 3, p. 7 |
| [4] | Greenhouse gas emissions from training (Llama 3.1 8B) | Llama 3.1 8B | 8B | Non spécifié | 420 tCO2e | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions' |
| [5] | Greenhouse gas emissions from training (Llama 3.1 70B) | Llama 3.1 70B | 70B | Non spécifié | 2040 tCO2e | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions' |
| [6] | Greenhouse gas emissions from training (Llama 3.1 405B) | Llama 3.1 405B | 405B | Non spécifié | 8930 tCO2e | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', table 'Training Location-Based Greenhouse Gas Emissions' |
| [7] | Compute time used for training (Llama 3.1 405B) | Llama 3.1 405B | 405B | Non spécifié | 30.84 million GPU-hours | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', cumulative compute table |
| [8] | Compute time used for training (Llama 3.1 8B) | Llama 3.1 8B | 8B | Non spécifié | 1.46 million GPU-hours | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', cumulative compute table |
| [9] | Compute time used for training (Llama 3.1 70B) | Llama 3.1 70B | 70B | Non spécifié | 7.0 million GPU-hours | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Hardware and Software', cumulative compute table |
| [10] | Training token volume (Llama 3.1 8B) | Llama 3.1 8B | 8B | Non spécifié | 15 trillion tokens | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Training Data' |
| [11] | Training token volume (Llama 3.1 70B) | Llama 3.1 70B | 70B | Non spécifié | 15 trillion tokens | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Training Data' |
| [12] | Training token volume (Llama 3.1 405B) | Llama 3.1 405B | 405B | Non spécifié | 15 trillion tokens | Meta (2024). Llama 3.1 Model Card. https://huggingface.co/meta-llama/Llama-3.1-405B Section 'Training Data' |
| [13] | Greenhouse gas emissions from training (Llama 4 Scout) | Llama 4 Scout | 109B | Non spécifié | 1354 tCO2e | Meta (2025) Llama 4 Model Card Section 'Hardware and Software', training emissions table |
| [14] | Greenhouse gas emissions from training (Llama 4 Maverick) | Llama 4 Maverick | 400B | Non spécifié | 645 tCO2e | Meta (2025) Llama 4 Model Card Section 'Hardware and Software', training emissions table |
| [15] | Training : training energy (Llama 4 Scout) | Llama 4 Scout | 109B | Non spécifié | 3.5 GWh | Meta (2025) Llama 4 Model Card Section 'Hardware and Software', 5.0M GPU hours on H100-80GB at 700W |
| [16] | Training : training energy (Llama 4 Maverick) | Llama 4 Maverick | 400B | Non spécifié | 1.666 GWh | Meta (2025) Llama 4 Model Card Section 'Hardware and Software', 2.38M GPU hours on H100-80GB at 700W |
| [17] | Training token volume (Llama 4 Scout) | Llama 4 Scout | 109B | Non spécifié | 40 trillion tokens | Meta (2025) Llama 4 Model Card Section 'Training Data' |
| [18] | Training token volume (Llama 4 Maverick) | Llama 4 Maverick | 400B | Non spécifié | 22 trillion tokens | Meta (2025) Llama 4 Model Card Section 'Training Data' |
| [19] | Emissions across the model creation lifecycle (OLMo 20M) | OLMo 20M | 20M | États-Unis | 0.3 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [20] | Emissions across the model creation lifecycle (OLMo 60M) | OLMo 60M | 60M | États-Unis | 0.4 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [21] | Emissions across the model creation lifecycle (OLMo 150M) | OLMo 150M | 150M | États-Unis | 1 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [22] | Emissions across the model creation lifecycle (OLMo 300M) | OLMo 300M | 300M | États-Unis | 2 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [23] | Emissions across the model creation lifecycle (OLMo 700M) | OLMo 700M | 700M | États-Unis | 3 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [24] | Emissions across the model creation lifecycle (OLMo 7B) | OLMo 7B | 7B | États-Unis | 22 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [25] | Emissions across the model creation lifecycle (OLMo 1B (3T)) | OLMo 1B (3T) | 1B | États-Unis | 10 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [26] | Emissions across the model creation lifecycle (OLMo 7B (Twin)) | OLMo 7B (Twin) | 7B | États-Unis | 70 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [27] | Emissions across the model creation lifecycle (OLMo (04|07)24 7B) | OLMo (04|07)24 7B | 7B | États-Unis | 32 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [28] | Emissions across the model creation lifecycle (OLMo 2 7B) | OLMo 2 7B | 7B | États-Unis | 52 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [29] | Emissions across the model creation lifecycle (OLMo 2 13B) | OLMo 2 13B | 13B | États-Unis | 101 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
| [30] | Emissions across the model creation lifecycle (OLMoE 0924) | OLMoE 0924 | 1B active / 7B total | États-Unis | 18 tCO2e | Morrison, J., Na, C., Fernandez, J., Dettmers, T., & others (2025). Holistically Evaluating the Environmental Impact of Creating Language Models. arXiv preprint arXiv:2503.05804. https://arxiv.org/abs/2503.05804 Table 2, p. 6; cluster locations in Section 3.1, p. 4 |
Real-world comparison benchmarks
| No. | Domain | Indicator | Value | Reference |
|---|---|---|---|---|
| 1 | Energy | Fluorescent lamp for 1 hour | 9,3 Wh | Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf Purdue Extension guide used for household-use examples; extraction still to be normalized more precisely. |
| 2 | Energy | Laptop for 1 hour | 32 Wh | Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf Purdue Extension guide used for household-use examples; extraction still to be normalized more precisely. |
| 3 | Energy | Laptop for 3 hours | 96 Wh | Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf Project derivation from the Purdue laptop example retained at 32 Wh for 1 hour, rescaled here to 3 hours of use. |
| 4 | Energy | Electric space heater for 1 hour | 1.5 kWh | Project calculation convention. Nominal power assumption fixed at 1,500 W, i.e. 1.5 kWh for 1 hour. |
| 5 | Energy | Laptop for 5 hours | 160 Wh | Carroll, N. J., & Kruse, J. (n.d.). Energy investigators 2: Facilitator’s guide. Purdue Extension. https://www.extension.purdue.edu/extmedia/4H/4-H-1015-W.pdf Project derivation from the Purdue laptop example retained at 32 Wh for 1 hour, rescaled here to 5 hours of use. |
| 6 | Energy | Electric kettle for 7 minutes | 210 Wh | Project calculation convention. Nominal power assumption fixed at 1,800 W, i.e. 210 Wh for 7 minutes of use. |
| 7 | Energy | Electric space heater for 10 minutes | 250 Wh | Project calculation convention. Project derivation from the retained 1,500 W space-heater convention, rescaled to 10 minutes of use. |
| 8 | Energy | 10,000 French households over one year of domestic use | 25 GWh | RTE. (2022, February 25). Bilan électrique 2021 - Une production d’électricité assurée à plus de 92% par des sources n’émettant pas de gaz à effet de serre. https://www.rte-france.com/actualites/bilan-electrique-2021 Project convention based on an average consumption of 2,500 kWh/year per household; comparison value made explicit in the training-chart note. |
| 9 | Carbon | Average gasoline car for 1 km | 235 gCO2e | International Council on Clean Transportation. (2025). Life-cycle greenhouse gas emissions from passenger cars in the European Union: A 2025 update and key factors to consider. https://theicct.org/publication/electric-cars-life-cycle-analysis-emissions-europe-jul25/ Key findings ; Figure 1 ; gasoline ICEV running on the average blend of fossil gasoline and ethanol estimated at 235 gCO2e/km, i.e. 235 gCO2e over 1 km. |
| 10 | Carbon | Average gasoline car for 430 m | 101.1 gCO2e | International Council on Clean Transportation. (2025). Life-cycle greenhouse gas emissions from passenger cars in the European Union: A 2025 update and key factors to consider. https://theicct.org/publication/electric-cars-life-cycle-analysis-emissions-europe-jul25/ Project derivation from the ICCT benchmark of 235 gCO2e/km for an average gasoline car, rescaled here to 430 meters to align with the current upper-end inference carbon tier retained in the comparison. |
| 11 | Carbon | Average full commercial flight (derived value) | ≈ 21.1 tCO2 per flight | Gössling, S., Klöwer, M., Leitão, J. C., Hirsch, S., Brockhagen, D., & Humpe, A. (2026). Large carbon dioxide emissions avoidance potential in improved commercial air transport efficiency. Communications Earth & Environment, 7, 13. https://www.nature.com/articles/s43247-025-03069-4 Results, “Emissions and efficiency”: 27,451,887 flights in 2023 causing 577,968,750 tCO2 emissions; the site then derives an average per flight. |
Central screening factors retained for market models
This table documents the central values retained by the project for the multi-factor prompt proxy of each quantified market model: raw active parameters, context window, serving mode, modality support, the resulting central factors F_ctx, F_srv, F_mod, F_arch, and the resulting central effective active-parameter proxy P_eff,c. The current annex covers 56 strict rows and 27 partial-data rows derived from donor models with sourced parameter counts.
| Model | Provider | Active parameters | Context window | Serving mode | Vision | F_ctx | F_srv | F_mod | F_arch | P_eff,c |
|---|---|---|---|---|---|---|---|---|---|---|
| ai21 | 178B [1] | 2,048 [2] | closed [4] | no [3] | 1.000 | 1.140 | 1.000 | 1.000 | 202.920B | |
| alibaba | 32B [7] | 131,072 [8] | open [10] | no [9] | 1.070 | 1.000 | 1.000 | 1.000 | 34.240B | |
| alibaba | 72B [13] | 131,072 [14] | open [16] | no [15] | 1.070 | 1.000 | 1.000 | 1.000 | 77.040B | |
| alibaba | 7B [19] | 131,072 [20] | open [22] | no [21] | 1.070 | 1.000 | 1.000 | 1.000 | 7.490B | |
| alibaba | 22B active / 235B total [25] | 131,072 [26] | open [28] | no [27] | 1.070 | 1.000 | 1.000 | 1.273 | 29.975B | |
| alibaba | 32.8B [31] | 131,072 [32] | open [34] | no [33] | 1.070 | 1.000 | 1.000 | 1.000 | 35.096B | |
| anthropic | 100B* [37] | 200,000 [38] | closed [40] | yes [39] | 1.091 | 1.140 | 1.030 | 1.000 | 128.145B | |
| anthropic | 100B* [43] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 128.145B | |
| anthropic | 100B* [43] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 128.145B | |
| anthropic | 100B* [43] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 128.145B | |
| anthropic | 22B* [48] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 28.192B | |
| anthropic | 175B* [37] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 224.253B | |
| anthropic | 175B* [37] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 224.253B | |
| anthropic | 5000B* [43] | 1,000,000 [44] | research [46] | yes [45] | 1.173 | 1.140 | 1.030 | 1.000 | 6884.363B | |
| anthropic | 2000B* [37] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 2562.897B | |
| anthropic | 5000B* [37] | 1,000,000 [44] | closed [46] | yes [45] | 1.173 | 1.140 | 1.030 | 1.000 | 6884.363B | |
| anthropic | 4000B* [50] | 1,000,000 [44] | closed [46] | yes [45] | 1.173 | 1.140 | 1.030 | 1.000 | 5507.491B | |
| anthropic | 175B* [43] | 200,000 [44] | closed [46] | yes [45] | 1.091 | 1.140 | 1.030 | 1.000 | 224.253B | |
| anthropic | 400B* [37] | 1,000,000 [44] | closed [46] | yes [45] | 1.173 | 1.140 | 1.030 | 1.000 | 550.749B | |
| deepmind | 280B [52] | 4,096 [53] | closed [55] | no [54] | 1.000 | 1.140 | 1.000 | 1.000 | 319.200B | |
| deepseek | 37B active / 671B total [58] | 128,000 [59] | hybrid [61] | no [60] | 1.069 | 1.070 | 1.000 | 1.401 | 59.289B | |
| deepseek | 37B active / 671B total [64] | 128,000 [65] | hybrid [67] | no [66] | 1.069 | 1.070 | 1.000 | 1.334 | 56.466B | |
| deepseek | 13B active / 284B total [70] | 1,000,000 [71] | hybrid [73] | no [72] | 1.173 | 1.070 | 1.000 | 1.356 | 22.117B | |
| deepseek | 49B active / 1600B total [76] | 1,000,000 [77] | hybrid [79] | no [78] | 1.173 | 1.070 | 1.000 | 1.472 | 90.526B | |
| 133.5B* [43] | 1,048,576 [82] | closed [84] | yes [83] | 1.175 | 1.140 | 1.030 | 1.000 | 184.188B | ||
| 133.5B* [43] | 1,048,576 [86] | closed [88] | yes [87] | 1.175 | 1.140 | 1.030 | 1.000 | 184.188B | ||
| 30B* [37] | 1,048,576 [90] | closed [92] | yes [91] | 1.175 | 1.140 | 1.030 | 1.000 | 41.391B | ||
| 80B* [94] | 1,048,576 [90] | closed [92] | yes [91] | 1.175 | 1.140 | 1.030 | 1.000 | 110.375B | ||
| 30B* [94] | 1,048,576 [90] | closed [92] | yes [91] | 1.175 | 1.140 | 1.030 | 1.000 | 41.391B | ||
| 200B* [37] | 1,048,576 [90] | closed [92] | yes [91] | 1.175 | 1.140 | 1.030 | 1.000 | 275.937B | ||
| 30B* [43] | 1,048,576 [96] | closed [98] | yes [97] | 1.175 | 1.140 | 1.030 | 1.000 | 41.391B | ||
| 3000B* [37] | 1,048,576 [100] | closed [102] | yes [101] | 1.175 | 1.140 | 1.030 | 1.000 | 4139.055B | ||
| 30B* [43] | 131,072 [104] | closed [106] | yes [105] | 1.070 | 1.140 | 1.030 | 1.000 | 37.692B | ||
| 22B* [43] | 8,192 [108] | closed [110] | no [109] | 1.000 | 1.140 | 1.000 | 1.000 | 25.080B | ||
| 22B* [43] | 1,048,576 [112] | closed [114] | yes [113] | 1.175 | 1.140 | 1.030 | 1.000 | 30.353B | ||
| 3000B* [37] | 1,048,576 [116] | closed [118] | yes [117] | 1.175 | 1.140 | 1.030 | 1.000 | 4139.055B | ||
| 130B [120] | 2,048 [121] | research [123] | no [122] | 1.000 | 1.140 | 1.000 | 1.000 | 148.200B | ||
| 137B* [126] | 2,048 [127] | closed [129] | no [128] | 1.000 | 1.140 | 1.000 | 1.000 | 156.180B | ||
| meta | 405B [132] | 131,072 [133] | open [135] | no [134] | 1.070 | 1.000 | 1.000 | 1.000 | 433.350B | |
| meta | 70B [132] | 131,072 [138] | open [140] | no [139] | 1.070 | 1.000 | 1.000 | 1.000 | 74.900B | |
| meta | 8B [132] | 131,072 [142] | open [144] | no [143] | 1.070 | 1.000 | 1.000 | 1.000 | 8.560B | |
| meta | 17B active / 400B total [146] | 1,000,000 [147] | open [149] | yes [148] | 1.173 | 1.000 | 1.030 | 1.365 | 28.017B | |
| meta | 17B active / 109B total [153] | 10,000,000 [154] | open [156] | yes [155] | 1.289 | 1.000 | 1.030 | 1.214 | 27.408B | |
| meta | 175B [160] | 2,048 [161] | open [163] | no [162] | 1.000 | 1.000 | 1.000 | 1.000 | 175.000B | |
| microsoft | 530B [166] | 2,048 [167] | closed [169] | no [168] | 1.000 | 1.140 | 1.000 | 1.000 | 604.200B | |
| mistral | 22B [172] | 32,000 [173] | hybrid [175] | no [174] | 1.000 | 1.070 | 1.000 | 1.000 | 23.540B | |
| mistral | 123B [178] | 256,000 [179] | hybrid [181] | no [180] | 1.104 | 1.070 | 1.000 | 1.000 | 145.271B | |
| mistral | 14B [184] | 256,000 [185] | hybrid [187] | yes [186] | 1.104 | 1.070 | 1.030 | 1.000 | 17.031B | |
| mistral | 3B [172] | 128,000 [190] | hybrid [192] | no [191] | 1.069 | 1.070 | 1.000 | 1.000 | 3.431B | |
| mistral | 8B [172] | 128,000 [194] | hybrid [196] | no [195] | 1.069 | 1.070 | 1.000 | 1.000 | 9.149B | |
| mistral | 123B [198] | 128,000 [199] | closed [201] | yes [200] | 1.069 | 1.140 | 1.030 | 1.000 | 154.364B | |
| mistral | 41B active / 675B total [204] | 256,000 [205] | hybrid [207] | yes [206] | 1.104 | 1.070 | 1.030 | 1.323 | 66.001B | |
| mistral | 128B* [210] | 256,000 [211] | hybrid [213] | yes [212] | 1.104 | 1.070 | 1.030 | 1.000 | 155.712B | |
| mistral | 24B [172] | 128,000 [216] | hybrid [218] | yes [217] | 1.069 | 1.070 | 1.030 | 1.000 | 28.270B | |
| mistral | 6.5B active / 119B total [220] | 256,000 [221] | hybrid [223] | yes [222] | 1.104 | 1.070 | 1.030 | 1.336 | 10.561B | |
| openai | 175B* [226] | 4,096 [227] | closed [229] | no [228] | 1.000 | 1.140 | 1.000 | 1.000 | 199.500B | |
| openai | 440B* active / 1760B* total [233] | 8,192 [234] | closed [236] | yes [235] | 1.000 | 1.140 | 1.030 | 1.160 | 599.312B | |
| openai | 200B* [50] | 128,000 [241] | closed [243] | yes [242] | 1.069 | 1.140 | 1.030 | 1.000 | 250.998B | |
| openai | 200B* [43] | 128,000 [245] | closed [247] | yes [246] | 1.069 | 1.140 | 1.030 | 1.000 | 250.998B | |
| openai | 95B* [249] | 400,000 [250] | closed [252] | yes [251] | 1.126 | 1.140 | 1.030 | 1.000 | 125.642B | |
| openai | 95B* [255] | 400,000 [256] | closed [258] | yes [257] | 1.126 | 1.140 | 1.030 | 1.000 | 125.642B | |
| openai | 3000B* [37] | 400,000 [261] | closed [263] | yes [262] | 1.126 | 1.140 | 1.030 | 1.000 | 3967.636B | |
| openai | 440B* [43] | 400,000 [265] | closed [267] | yes [266] | 1.126 | 1.140 | 1.030 | 1.000 | 581.920B | |
| openai | 3000B* [37] | 400,000 [269] | closed [271] | yes [270] | 1.126 | 1.140 | 1.030 | 1.000 | 3967.636B | |
| openai | 95B* [43] | 400,000 [273] | closed [275] | yes [274] | 1.126 | 1.140 | 1.030 | 1.000 | 125.642B | |
| openai | 95B* [43] | 400,000 [277] | closed [279] | yes [278] | 1.126 | 1.140 | 1.030 | 1.000 | 125.642B | |
| openai | 3000B* [43] | 400,000 [281] | closed [283] | yes [282] | 1.126 | 1.140 | 1.030 | 1.000 | 3967.636B | |
| openai | 3000B* [43] | 400,000 [285] | closed [287] | yes [286] | 1.126 | 1.140 | 1.030 | 1.000 | 3967.636B | |
| openai | 3000B* [43] | 1,050,000 [289] | closed [291] | yes [290] | 1.175 | 1.140 | 1.030 | 1.000 | 4139.296B | |
| openai | 5.1B active / 117B total [293] | 131,072 [294] | open [296] | no [295] | 1.070 | 1.000 | 1.000 | 1.362 | 7.430B | |
| openai | 3.6B active / 21B total [299] | 131,072 [300] | open [302] | no [301] | 1.070 | 1.000 | 1.000 | 1.204 | 4.636B | |
| openai | 320B* [43] | 200,000 [305] | closed [307] | yes [306] | 1.091 | 1.140 | 1.030 | 1.000 | 410.063B | |
| openai | 200B* [43] | 128,000 [309] | closed [311] | no [310] | 1.069 | 1.140 | 1.000 | 1.050 | 255.871B | |
| openai | 320B* [43] | 128,000 [313] | closed [315] | no [314] | 1.069 | 1.140 | 1.000 | 1.050 | 409.394B | |
| xai | 78.5B* active / 314B* total [317] | 200,000 [318] | closed [320] | yes [319] | 1.091 | 1.140 | 1.030 | 1.160 | 116.689B | |
| xai | 78.5B* [43] | 128,000 [324] | closed [326] | no [325] | 1.069 | 1.140 | 1.000 | 1.050 | 100.429B | |
| xai | 115B* active / 270B* total [328] | 2,000,000 [318] | closed [320] | yes [319] | 1.208 | 1.140 | 1.030 | 1.099 | 179.130B | |
| xai | 600B* [37] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.050 | 893.321B | |
| xai | 96.75B* [43] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.000 | 137.189B | |
| xai | 96.75B* [43] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.050 | 144.048B | |
| xai | 600B* [43] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.000 | 850.782B | |
| xai | 600B* [43] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.050 | 893.321B | |
| xai | 4000B* [43] | 2,000,000 [324] | closed [326] | yes [325] | 1.208 | 1.140 | 1.030 | 1.000 | 5671.879B |
Central training screening factors retained for market models
This table documents the central values retained by the project for the multi-factor training proxy of each quantified market model: retained training parameter count, training-token prior, training regime, multimodal training flag, hardware-class proxy, and the resulting central factors F_reg, F_arch-tr, and F_hw. These are project screening factors, not provider-published measurements. When a field is retained as an estimate or screening prior, its citation anchors the release line or methodological basis rather than an exact provider-published numeric value.
| Model | Provider | Retained parameters | Training tokens | Training regime | Multimodal | Hardware class | F_reg | F_arch | F_hw |
|---|---|---|---|---|---|---|---|---|---|
| ai21 | 178B [5] | 5.00T* | pretraining* | no [6] | standard_gpu_cluster | 1.0000 | 1.0000 | 1.0500 | |
| alibaba | 32B [11] | 0.64T* | pretraining* | no [12] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| alibaba | 72B [17] | 1.44T* | pretraining* | no [18] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| alibaba | 7B [23] | 0.14T* | pretraining* | no [24] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| alibaba | 235B [29] | 4.70T* | pretraining* | no [30] | standard_gpu_cluster* | 1.0000 | 0.9000 | 1.0500 | |
| alibaba | 32.8B [35] | 0.66T* | pretraining* | no [36] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| anthropic | 100B [41] | 2.00T* | pretraining* | yes [42] | modern_hyperscale_gpu | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 100B* | 0.44T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 100B* | 5.20T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 100B* | 2.40T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 22B [49] | 0.44T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 175B [41] | 3.50T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 175B [41] | 3.50T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 5000B* | 52.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 2000B [41] | 40.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 5000B [41] | 100.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 4000B [51] | 80.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 175B* | 8.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| anthropic | 400B [41] | 8.00T* | pretraining* | yes [47] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| deepmind | 280B [56] | 6.00T* | pretraining* | no [57] | modern_hyperscale_gpu | 1.0000 | 1.0000 | 0.9000 | |
| deepseek | 671B [62] | 13.42T* | pretraining* | no [63] | mixed_gpu_cluster* | 1.0000 | 0.9000 | 1.0000 | |
| deepseek | 671B [68] | 13.42T* | pretraining* | no [69] | mixed_gpu_cluster* | 1.0000 | 0.9000 | 1.0000 | |
| deepseek | 284B [74] | 5.68T* | pretraining* | no [75] | mixed_gpu_cluster* | 1.0000 | 0.9000 | 1.0000 | |
| deepseek | 1600B [80] | 32.00T* | pretraining* | no [81] | mixed_gpu_cluster* | 1.0000 | 0.9000 | 1.0000 | |
| 133.5B* | 1.20T* | pretraining* | yes [85] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 133.5B* | 2.40T* | pretraining* | yes [89] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 30B [41] | 0.60T* | pretraining* | yes [93] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 80B [95] | 1.60T* | pretraining* | yes [93] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 30B [95] | 0.60T* | pretraining* | yes [93] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 200B [41] | 4.00T* | pretraining* | yes [93] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 30B* | 2.00T* | pretraining* | yes [99] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 3000B [41] | 60.00T* | pretraining* | yes [103] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 30B* | 2.00T* | pretraining* | yes [107] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 22B* | 0.80T* | pretraining* | yes [111] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 22B* | 0.80T* | pretraining* | yes [115] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 3000B [41] | 60.00T* | pretraining* | yes [119] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | ||
| 130B [124] | 4.00T* | pretraining* | no [125] | modern_hyperscale_gpu | 1.0000 | 0.9000 | 0.9000 | ||
| 137B [130] | 3.00T* | pretraining* | no [131] | modern_hyperscale_gpu | 1.0000 | 1.0000 | 0.9000 | ||
| meta | 405B [136] | 8.10T* | pretraining* | no [137] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| meta | 70B [136] | 1.40T* | pretraining* | no [141] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| meta | 8B [136] | 0.16T* | pretraining* | no [145] | standard_gpu_cluster* | 1.0000 | 1.0000 | 1.0500 | |
| meta | 400B [150] | 22.00T [151] | pretraining* | yes [152] | standard_gpu_cluster* | 1.0000 | 1.0350 | 1.0500 | |
| meta | 109B [157] | 40.00T [158] | pretraining* | yes [159] | standard_gpu_cluster* | 1.0000 | 1.0350 | 1.0500 | |
| meta | 175B [164] | 4.00T* | pretraining* | no [165] | standard_gpu_cluster | 1.0000 | 1.0000 | 1.0500 | |
| microsoft | 530B [170] | 20.00T* | pretraining* | no [171] | modern_hyperscale_gpu | 1.0000 | 1.0000 | 0.9000 | |
| mistral | 22B [176] | 0.44T* | pretraining* | no [177] | mixed_gpu_cluster* | 1.0000 | 1.0000 | 1.0000 | |
| mistral | 123B [182] | 2.46T* | pretraining* | no [183] | mixed_gpu_cluster* | 1.0000 | 1.0000 | 1.0000 | |
| mistral | 14B [188] | 0.28T* | pretraining* | yes [189] | mixed_gpu_cluster* | 1.0000 | 1.1500 | 1.0000 | |
| mistral | 3B [176] | 0.06T* | pretraining* | no [193] | mixed_gpu_cluster* | 1.0000 | 1.0000 | 1.0000 | |
| mistral | 8B [176] | 0.16T* | pretraining* | no [197] | mixed_gpu_cluster* | 1.0000 | 1.0000 | 1.0000 | |
| mistral | 123B [202] | 2.46T* | pretraining* | yes [203] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| mistral | 675B [208] | 13.50T* | pretraining* | yes [209] | mixed_gpu_cluster* | 1.0000 | 1.0350 | 1.0000 | |
| mistral | 128B [214] | 2.56T* | pretraining* | yes [215] | mixed_gpu_cluster* | 1.0000 | 1.1500 | 1.0000 | |
| mistral | 24B [176] | 0.48T* | pretraining* | yes [219] | mixed_gpu_cluster* | 1.0000 | 1.1500 | 1.0000 | |
| mistral | 119B [224] | 2.38T* | pretraining* | yes [225] | mixed_gpu_cluster* | 1.0000 | 1.0350 | 1.0000 | |
| openai | 175B [230] | 3.50T* | pretraining* | no [231] | standard_gpu_cluster [232] | 1.0000 | 1.0000 | 1.0500 | |
| openai | 1760B [237] | 6.00T* | pretraining [238] | yes [239] | modern_hyperscale_gpu [240] | 1.0000 | 1.0350 | 0.9000 | |
| openai | 200B [51] | 4.00T* | pretraining* | yes [244] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 200B* | 0.16T* | pretraining* | yes [248] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 95B [253] | 1.90T* | pretraining* | yes [254] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 95B [259] | 1.90T* | pretraining* | yes [260] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 3000B [41] | 60.00T* | pretraining* | yes [264] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 440B* | 6.00T* | pretraining* | yes [268] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 3000B [41] | 60.00T* | pretraining* | yes [272] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 95B* | 1.90T* | pretraining* | yes [276] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 95B* | 1.90T* | pretraining* | yes [280] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 3000B* | 36.00T* | pretraining* | yes [284] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 3000B* | 40.00T* | pretraining* | yes [288] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 3000B* | 40.00T* | pretraining* | yes [292] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 117B [297] | 2.34T* | pretraining* | no [298] | standard_gpu_cluster* | 1.0000 | 0.9000 | 1.0500 | |
| openai | 21B [303] | 0.42T* | pretraining* | no [304] | standard_gpu_cluster* | 1.0000 | 0.9000 | 1.0500 | |
| openai | 320B* | 6.40T* | pretraining* | yes [308] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| openai | 200B* | 1.90T* | pretraining* | no [312] | modern_hyperscale_gpu* | 1.0000 | 1.0000 | 0.9000 | |
| openai | 320B* | 6.00T* | pretraining* | no [316] | modern_hyperscale_gpu* | 1.0000 | 1.0000 | 0.9000 | |
| xai | 314B [321] | 6.28T* | pretraining [322] | yes [323] | modern_hyperscale_gpu | 1.0000 | 1.0350 | 0.9000 | |
| xai | 78.5B* | 1.80T* | pretraining* | no [327] | modern_hyperscale_gpu* | 1.0000 | 1.0000 | 0.9000 | |
| xai | 270B [329] | 5.40T* | pretraining* | yes [323] | modern_hyperscale_gpu | 1.0000 | 1.0350 | 0.9000 | |
| xai | 600B [41] | 12.00T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| xai | 96.75B* | 2.40T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| xai | 96.75B* | 2.40T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| xai | 600B* | 13.00T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| xai | 600B* | 13.00T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 | |
| xai | 4000B* | 13.00T* | pretraining* | yes [327] | modern_hyperscale_gpu* | 1.0000 | 1.1500 | 0.9000 |
* indicates a value retained from the project's screening rather than directly linked to an external source in this table.
Numbered source list for retained screening characteristics
The numbered references used in the retained inference and training screening-characteristic tables are listed below.
- [1] Active parameters. AI21 Jurassic-1 announcement
- [2] Context window. AI21 admissions
- [3] Vision. AI21 blog
- [4] Serving mode. AI21 admissions
- [5] Training parameters. AI21 Jurassic-1 announcement
- [6] Training modality. Jurassic-1 platform blog describing text API
- [7] Active parameters. Qwen2.5 32B model card
- [8] Context window. Qwen2.5 32B model card
- [9] Vision. Qwen2.5 32B model card
- [10] Serving mode. Qwen2.5 32B model card
- [11] Training parameters. Qwen2.5 32B model card
- [12] Training modality. Qwen2.5 32B model card
- [13] Active parameters. Qwen2.5 72B model card
- [14] Context window. Qwen2.5 72B model card
- [15] Vision. Qwen2.5 72B model card
- [16] Serving mode. Qwen2.5 72B model card
- [17] Training parameters. Qwen2.5 72B model card
- [18] Training modality. Qwen2.5 72B model card
- [19] Active parameters. Qwen2.5 7B model card
- [20] Context window. Qwen2.5 7B model card
- [21] Vision. Qwen2.5 7B model card
- [22] Serving mode. Qwen2.5 7B model card
- [23] Training parameters. Qwen2.5 7B model card
- [24] Training modality. Qwen2.5 7B model card
- [25] Active parameters. Qwen3 235B A22B release repository
- [26] Context window. Qwen3 235B A22B release repository
- [27] Vision. Qwen3 235B A22B release repository
- [28] Serving mode. Qwen3 235B A22B release repository
- [29] Training parameters. Qwen3 235B A22B release repository
- [30] Training modality. Qwen3 235B A22B release repository
- [31] Active parameters. Qwen3 32B release repository
- [32] Context window. Qwen3 32B release repository
- [33] Vision. Qwen3 32B release repository
- [34] Serving mode. Qwen3 32B release repository
- [35] Training parameters. Qwen3 32B release repository
- [36] Training modality. Qwen3 32B release repository
- [37] Active parameters. Third-party screening estimate from Alan D. Thompson Models Table
- [38] Context window. Anthropic Claude overview
- [39] Vision. Anthropic Claude overview
- [40] Serving mode. Anthropic Claude overview
- [41] Training parameters. Third-party screening estimate from Alan D. Thompson Models Table
- [42] Training modality. https://docs.anthropic.com/en/docs/about-claude/models/overview
- [43] Active parameters. Partial-data donor prior from source-linked market models
- [44] Context window. Models overview - Claude API Docs
- [45] Vision. Models overview - Claude API Docs
- [46] Serving mode. Models overview - Claude API Docs
- [47] Training modality. Models overview - Claude API Docs
- [48] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [49] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [50] Active parameters. Third-party screening estimate from Alan D. Thompson Models Table
- [51] Training parameters. Third-party screening estimate from Alan D. Thompson Models Table
- [52] Active parameters. DeepMind Gopher paper
- [53] Context window. DeepMind blog
- [54] Vision. DeepMind blog
- [55] Serving mode. DeepMind blog
- [56] Training parameters. DeepMind Gopher paper
- [57] Training modality. DeepMind Gopher paper centered on text modeling
- [58] Active parameters. DeepSeek-R1 model card
- [59] Context window. deepseek-ai/DeepSeek-R1
- [60] Vision. deepseek-ai/DeepSeek-R1
- [61] Serving mode. deepseek-ai/DeepSeek-R1
- [62] Training parameters. DeepSeek-R1 model card
- [63] Training modality. deepseek-ai/DeepSeek-R1
- [64] Active parameters. DeepSeek-V3 model card
- [65] Context window. deepseek-ai/DeepSeek-V3
- [66] Vision. deepseek-ai/DeepSeek-V3
- [67] Serving mode. deepseek-ai/DeepSeek-V3
- [68] Training parameters. DeepSeek-V3 model card
- [69] Training modality. deepseek-ai/DeepSeek-V3
- [70] Active parameters. DeepSeek V4 Flash official release note
- [71] Context window. DeepSeek V4 Flash documentation | DeepSeek
- [72] Vision. DeepSeek V4 Flash documentation | DeepSeek
- [73] Serving mode. DeepSeek V4 Flash documentation | DeepSeek
- [74] Training parameters. DeepSeek V4 Flash official release note
- [75] Training modality. DeepSeek V4 Flash documentation | DeepSeek
- [76] Active parameters. DeepSeek V4 Pro official release note
- [77] Context window. DeepSeek V4 Pro documentation | DeepSeek
- [78] Vision. DeepSeek V4 Pro documentation | DeepSeek
- [79] Serving mode. DeepSeek V4 Pro documentation | DeepSeek
- [80] Training parameters. DeepSeek V4 Pro official release note
- [81] Training modality. DeepSeek V4 Pro documentation | DeepSeek
- [82] Context window. Gemini 1.5 Flash | Gemini API
- [83] Vision. Gemini 1.5 Flash | Gemini API
- [84] Serving mode. Gemini 1.5 Flash | Gemini API
- [85] Training modality. Gemini 1.5 Flash | Gemini API
- [86] Context window. Gemini 1.5 Pro | Gemini API
- [87] Vision. Gemini 1.5 Pro | Gemini API
- [88] Serving mode. Gemini 1.5 Pro | Gemini API
- [89] Training modality. Gemini 1.5 Pro | Gemini API
- [90] Context window. Gemini models | Gemini API
- [91] Vision. Gemini models | Gemini API
- [92] Serving mode. Gemini models | Gemini API
- [93] Training modality. Gemini models | Gemini API
- [94] Active parameters. Third-party family-level screening estimate from Alan D. Thompson Models Table
- [95] Training parameters. Third-party family-level screening estimate from Alan D. Thompson Models Table
- [96] Context window. Gemini 3 Flash Preview | Gemini API
- [97] Vision. Gemini 3 Flash Preview | Gemini API
- [98] Serving mode. Gemini 3 Flash Preview | Gemini API
- [99] Training modality. Gemini 3 Flash Preview | Gemini API
- [100] Context window. Gemini 3 Pro Preview | Gemini API
- [101] Vision. Gemini 3 Pro Preview | Gemini API
- [102] Serving mode. Gemini 3 Pro Preview | Gemini API
- [103] Training modality. Gemini 3 Pro Preview | Gemini API
- [104] Context window. Gemini 3.1 Flash Live Preview | Gemini API
- [105] Vision. Gemini 3.1 Flash Live Preview | Gemini API
- [106] Serving mode. Gemini 3.1 Flash Live Preview | Gemini API
- [107] Training modality. Gemini 3.1 Flash Live Preview | Gemini API
- [108] Context window. Gemini 3.1 Flash TTS Preview | Gemini API
- [109] Vision. Gemini 3.1 Flash TTS Preview | Gemini API
- [110] Serving mode. Gemini 3.1 Flash TTS Preview | Gemini API
- [111] Training modality. Gemini 3.1 Flash TTS Preview | Gemini API
- [112] Context window. Gemini 3.1 Flash-Lite Preview | Gemini API
- [113] Vision. Gemini 3.1 Flash-Lite Preview | Gemini API
- [114] Serving mode. Gemini 3.1 Flash-Lite Preview | Gemini API
- [115] Training modality. Gemini 3.1 Flash-Lite Preview | Gemini API
- [116] Context window. Gemini 3.1 Pro Preview | Gemini API
- [117] Vision. Gemini 3.1 Pro Preview | Gemini API
- [118] Serving mode. Gemini 3.1 Pro Preview | Gemini API
- [119] Training modality. Gemini 3.1 Pro Preview | Gemini API
- [120] Active parameters. Google AI blog
- [121] Context window. Google AI blog
- [122] Vision. Google AI blog
- [123] Serving mode. Google AI blog
- [124] Training parameters. Google AI blog
- [125] Training modality. GLaM blog details conditional computation for text
- [126] Active parameters. LaMDA research paper / Google Research summary
- [127] Context window. Google blog
- [128] Vision. Google blog
- [129] Serving mode. Google blog
- [130] Training parameters. LaMDA research paper / Google Research summary
- [131] Training modality. LaMDA announcement describing dialog-only model
- [132] Active parameters. Meta Llama 3.1 model card
- [133] Context window. meta-llama/Llama-3.1 model card
- [134] Vision. meta-llama/Llama-3.1 model card
- [135] Serving mode. meta-llama/Llama-3.1 model card
- [136] Training parameters. Meta Llama 3.1 model card
- [137] Training modality. meta-llama/Llama-3.1 model card
- [138] Context window. meta-llama/Llama-3.1 model card
- [139] Vision. meta-llama/Llama-3.1 model card
- [140] Serving mode. meta-llama/Llama-3.1 model card
- [141] Training modality. meta-llama/Llama-3.1 model card
- [142] Context window. meta-llama/Llama-3.1 model card
- [143] Vision. meta-llama/Llama-3.1 model card
- [144] Serving mode. meta-llama/Llama-3.1 model card
- [145] Training modality. meta-llama/Llama-3.1 model card
- [146] Active parameters. Llama 4 Maverick model card
- [147] Context window. Llama 4 Maverick model card
- [148] Vision. Llama 4 Maverick model card
- [149] Serving mode. Llama 4 Maverick model card
- [150] Training parameters. Llama 4 Maverick model card
- [151] Training tokens. Official model card training-token disclosure (22T tokens)
- [152] Training modality. Llama 4 Maverick model card
- [153] Active parameters. Llama 4 Scout model card
- [154] Context window. Llama 4 Scout model card
- [155] Vision. Llama 4 Scout model card
- [156] Serving mode. Llama 4 Scout model card
- [157] Training parameters. Llama 4 Scout model card
- [158] Training tokens. Official model card training-token disclosure (40T tokens)
- [159] Training modality. Llama 4 Scout model card
- [160] Active parameters. OPT release
- [161] Context window. Meta AI blog
- [162] Vision. Meta AI blog
- [163] Serving mode. Meta AI blog
- [164] Training parameters. OPT release
- [165] Training modality. OPT release blog describing dense decoder transformer
- [166] Active parameters. NVidia & Microsoft announcement
- [167] Context window. Microsoft product page
- [168] Vision. Microsoft announcement
- [169] Serving mode. Microsoft product page
- [170] Training parameters. NVidia & Microsoft announcement
- [171] Training modality. Azure AI service introduction describes a text-only transformer
- [172] Active parameters. Mistral model overview
- [173] Context window. Codestral - Mistral Docs
- [174] Vision. Codestral - Mistral Docs
- [175] Serving mode. Codestral - Mistral Docs
- [176] Training parameters. Mistral model overview
- [177] Training modality. Codestral - Mistral Docs
- [178] Active parameters. Devstral 2 model documentation
- [179] Context window. Devstral 2 model documentation | Mistral
- [180] Vision. Devstral 2 model documentation | Mistral
- [181] Serving mode. Devstral 2 model documentation | Mistral
- [182] Training parameters. Devstral 2 model documentation
- [183] Training modality. Devstral 2 model documentation | Mistral
- [184] Active parameters. Ministral 3 14B model documentation
- [185] Context window. Ministral 3 14B model documentation | Mistral
- [186] Vision. Ministral 3 14B model documentation | Mistral
- [187] Serving mode. Ministral 3 14B model documentation | Mistral
- [188] Training parameters. Ministral 3 14B model documentation
- [189] Training modality. Ministral 3 14B model documentation | Mistral
- [190] Context window. Research-licensed edge model - Mistral Docs
- [191] Vision. Research-licensed edge model - Mistral Docs
- [192] Serving mode. Research-licensed edge model - Mistral Docs
- [193] Training modality. Research-licensed edge model - Mistral Docs
- [194] Context window. Research-licensed edge model - Mistral Docs
- [195] Vision. Research-licensed edge model - Mistral Docs
- [196] Serving mode. Research-licensed edge model - Mistral Docs
- [197] Training modality. Research-licensed edge model - Mistral Docs
- [198] Active parameters. Mistral Large 2 launch note
- [199] Context window. Mistral Large 2.1 - Mistral Docs
- [200] Vision. Mistral Large 2.1 - Mistral Docs
- [201] Serving mode. Mistral Large 2.1 - Mistral Docs
- [202] Training parameters. Mistral Large 2 launch note
- [203] Training modality. Mistral Large 2.1 - Mistral Docs
- [204] Active parameters. Mistral Large 3 model documentation
- [205] Context window. Mistral Large 3 model documentation | Mistral
- [206] Vision. Mistral Large 3 model documentation | Mistral
- [207] Serving mode. Mistral Large 3 model documentation | Mistral
- [208] Training parameters. Mistral Large 3 model documentation
- [209] Training modality. Mistral Large 3 model documentation | Mistral
- [210] Active parameters. Third-party screening estimate from Artificial Analysis
- [211] Context window. Mistral Medium 3.5 model documentation | Mistral
- [212] Vision. Mistral Medium 3.5 model documentation | Mistral
- [213] Serving mode. Mistral Medium 3.5 model documentation | Mistral
- [214] Training parameters. Third-party screening estimate from Artificial Analysis
- [215] Training modality. Mistral Medium 3.5 model documentation | Mistral
- [216] Context window. Mistral Small 3.1 - Mistral Docs
- [217] Vision. Mistral Small 3.1 - Mistral Docs
- [218] Serving mode. Mistral Small 3.1 - Mistral Docs
- [219] Training modality. Mistral Small 3.1 - Mistral Docs
- [220] Active parameters. Mistral Small 4 model documentation
- [221] Context window. Mistral Small 4 model documentation | Mistral
- [222] Vision. Mistral Small 4 model documentation | Mistral
- [223] Serving mode. Mistral Small 4 model documentation | Mistral
- [224] Training parameters. Mistral Small 4 model documentation
- [225] Training modality. Mistral Small 4 model documentation | Mistral
- [226] Active parameters. Third-party screening estimate from Nexos
- [227] Context window. OpenAI ChatGPT blog
- [228] Vision. OpenAI ChatGPT blog
- [229] Serving mode. OpenAI ChatGPT blog
- [230] Training parameters. Third-party screening estimate from Nexos
- [231] Training modality. OpenAI ChatGPT blog
- [232] Training hardware. OpenAI ChatGPT blog
- [233] Active parameters. Third-party parameter estimate from Exploding Topics
- [234] Context window. OpenAI GPT-4 research release
- [235] Vision. OpenAI GPT-4 research release
- [236] Serving mode. OpenAI GPT-4 research release
- [237] Training parameters. Third-party parameter estimate from Exploding Topics
- [238] Training regime. OpenAI GPT-4 technical report
- [239] Training modality. OpenAI GPT-4 research release
- [240] Training hardware. OpenAI GPT-4 research release
- [241] Context window. GPT-4o model docs | OpenAI API
- [242] Vision. GPT-4o model docs | OpenAI API
- [243] Serving mode. GPT-4o model docs | OpenAI API
- [244] Training modality. GPT-4o model docs | OpenAI API
- [245] Context window. GPT-4o mini model docs | OpenAI API
- [246] Vision. Introducing GPT-4o mini | OpenAI
- [247] Serving mode. GPT-4o mini model docs | OpenAI API
- [248] Training modality. Introducing GPT-4o mini | OpenAI
- [249] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [250] Context window. GPT-5 mini Model | OpenAI API
- [251] Vision. GPT-5 mini Model | OpenAI API
- [252] Serving mode. GPT-5 mini Model | OpenAI API
- [253] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [254] Training modality. GPT-5 mini Model | OpenAI API
- [255] Active parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [256] Context window. GPT-5 nano Model | OpenAI API
- [257] Vision. GPT-5 nano Model | OpenAI API
- [258] Serving mode. GPT-5 nano Model | OpenAI API
- [259] Training parameters. Artificial Analysis size-class midpoint retained as screening estimate
- [260] Training modality. GPT-5 nano Model | OpenAI API
- [261] Context window. GPT-5.2 model docs | OpenAI API
- [262] Vision. GPT-5.2 model docs | OpenAI API
- [263] Serving mode. GPT-5.2 model docs | OpenAI API
- [264] Training modality. GPT-5.2 model docs | OpenAI API
- [265] Context window. GPT-5.2-pro model docs | OpenAI API
- [266] Vision. GPT-5.2-pro model docs | OpenAI API
- [267] Serving mode. GPT-5.2-pro model docs | OpenAI API
- [268] Training modality. GPT-5.2-pro model docs | OpenAI API
- [269] Context window. GPT-5.4 model docs | OpenAI API
- [270] Vision. GPT-5.4 model docs | OpenAI API
- [271] Serving mode. GPT-5.4 model docs | OpenAI API
- [272] Training modality. GPT-5.4 model docs | OpenAI API
- [273] Context window. GPT-5.4 mini model docs | OpenAI API
- [274] Vision. GPT-5.4 mini model docs | OpenAI API
- [275] Serving mode. GPT-5.4 mini model docs | OpenAI API
- [276] Training modality. GPT-5.4 mini model docs | OpenAI API
- [277] Context window. GPT-5.4 nano model docs | OpenAI API
- [278] Vision. GPT-5.4 nano model docs | OpenAI API
- [279] Serving mode. GPT-5.4 nano model docs | OpenAI API
- [280] Training modality. GPT-5.4 nano model docs | OpenAI API
- [281] Context window. GPT-5.4-pro model docs | OpenAI API
- [282] Vision. GPT-5.4-pro model docs | OpenAI API
- [283] Serving mode. GPT-5.4-pro model docs | OpenAI API
- [284] Training modality. GPT-5.4-pro model docs | OpenAI API
- [285] Context window. GPT-5.5 model docs | OpenAI API
- [286] Vision. GPT-5.5 model docs | OpenAI API
- [287] Serving mode. GPT-5.5 model docs | OpenAI API
- [288] Training modality. GPT-5.5 model docs | OpenAI API
- [289] Context window. GPT-5.5-pro model docs | OpenAI API
- [290] Vision. GPT-5.5-pro model docs | OpenAI API
- [291] Serving mode. GPT-5.5-pro model docs | OpenAI API
- [292] Training modality. GPT-5.5-pro model docs | OpenAI API
- [293] Active parameters. gpt-oss-120b Model | OpenAI API
- [294] Context window. gpt-oss-120b Model | OpenAI API
- [295] Vision. gpt-oss-120b Model | OpenAI API
- [296] Serving mode. gpt-oss-120b Model | OpenAI API
- [297] Training parameters. gpt-oss-120b Model | OpenAI API
- [298] Training modality. gpt-oss-120b Model | OpenAI API
- [299] Active parameters. gpt-oss-20b Model | OpenAI API
- [300] Context window. gpt-oss-20b Model | OpenAI API
- [301] Vision. gpt-oss-20b Model | OpenAI API
- [302] Serving mode. gpt-oss-20b Model | OpenAI API
- [303] Training parameters. gpt-oss-20b Model | OpenAI API
- [304] Training modality. gpt-oss-20b Model | OpenAI API
- [305] Context window. o1 model docs | OpenAI API
- [306] Vision. o1 model docs | OpenAI API
- [307] Serving mode. o1 model docs | OpenAI API
- [308] Training modality. o1 model docs | OpenAI API
- [309] Context window. o1-mini model docs | OpenAI API
- [310] Vision. o1-mini model docs | OpenAI API
- [311] Serving mode. o1-mini model docs | OpenAI API
- [312] Training modality. o1-mini model docs | OpenAI API
- [313] Context window. o1-preview model docs | OpenAI API
- [314] Vision. o1-preview model docs | OpenAI API
- [315] Serving mode. o1-preview model docs | OpenAI API
- [316] Training modality. o1-preview model docs | OpenAI API
- [317] Active parameters. Open Release of Grok-1 | xAI
- [318] Context window. https://docs.x.ai/developers/models
- [319] Vision. https://docs.x.ai/developers/models
- [320] Serving mode. https://docs.x.ai/developers/models
- [321] Training parameters. Open Release of Grok-1 | xAI
- [322] Training regime. Open Release of Grok-1 | xAI
- [323] Training modality. https://docs.x.ai/developers/models
- [324] Context window. Models and Pricing | xAI
- [325] Vision. Models and Pricing | xAI
- [326] Serving mode. Models and Pricing | xAI
- [327] Training modality. Models and Pricing | xAI
- [328] Active parameters. Third-party screening estimate from the Grok-2 open-weight config discussion
- [329] Training parameters. Third-party screening estimate from the Grok-2 open-weight config discussion
Country factors for carbon and water recalculation
| No. | Country | Year | Carbon intensity | Reference |
|---|---|---|---|---|
| 1 | France | 2024 | 40 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |
| 2 | United States | 2024 | 385 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |
| 3 | Germany | 2024 | 380 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |
| 4 | United Kingdom | 2024 | 180 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |
| 5 | Canada | 2024 | 120 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |
| 6 | China | 2024 | 540 gCO2e/kWh | ImpactLLM method note: project screening electricity factors Project screening default retained for rapid comparative estimation; not an audited national inventory factor. |