How the index is calculated
The index combines model-level evaluation ranks with company-level governance evidence. The overall ranking is a 90% behavior-rank / 10% governance-percentile blend. Behavior uses the mean ordinal rank across complete-lineage bootstrap refits with the taxonomy-v3 component budgets and quality weights.
1. Collect and group evaluations, and standardize model identities
The current version analyzes 817 models using 111 evals and 303 sub-evals from 119 source datasets. Related score columns are grouped under one eval so that a benchmark does not gain weight by reporting more columns.
Reasoning settings such as low, medium, high, or “thinking” are treated as configurations of the same model. If several settings appear in one sub-eval, their scores are averaged.
Dated API snapshots from the same product generation are also pooled. If one sub-eval reports several dated snapshots, the newest verified checkpoint is used; ties are averaged. Undated aliases lose to a dated snapshot when both appear together. Model sizes, product generations, fine-tunes, quantizations, safety settings, and multi-agent systems remain separate.
2. Fit the ranking
Each sub-eval is oriented so that better results rank first. Historical complete-lineage bootstrap draws refit the fixed-component weighted-Spearman ordinal estimator; their mean rank defines the public behavior ordering. A separate bootstrap over transitive governed-source dependency components supplies corpus-source sensitivity intervals. Native score magnitudes never cross eval boundaries.
The public evidence threshold is applied only after fitting. Lower rank numbers are better; there is no latent cardinal behavior score.
3. Calculate behavior rank sensitivities
The seven-component taxonomy assigns 25% to nonhuman ethics, 15% to human rights and systemic harm, 10% to fairness, 15% to truthfulness and honesty, 10% to benign helpfulness, 10% to ordinary-harm resistance, and 15% to responsible agency. For each model and component, the component table removes only that model's observations assigned to the component, refits the ordering, and reports the model's change in rank. These leave-one-model-component-out effects are sensitivities and do not sum.
Each eval receives one vote before component weights are applied. A bounded quality multiplier accounts for construct relevance, breadth, provenance, and overlap with other evals. Knowledge, recognition, classification, calibration, constrained-QA, and stated-tendency measures are admissible when mapped to the component they actually measure; protocol limitations are disclosed and reflected in lineage weights.
4. Add company governance
Governance contributes 10% and comes from the Future of Life Institute AI Safety Index. It averages five published domains: risk assessment, safety frameworks, existential safety, governance and accountability, and information sharing. The Current Harms domain is excluded because behavioral harms are already represented in the evals.
For each model, the developer's nearest available index edition is matched to the model release date. Downstream models may inherit the base developer's score when the relationship is explicit. When the index covers neither the developer nor a mapped base developer, the governance component uses the median across the available company-edition scores. The governance value and behavior rank are each converted to percentile utility before the 90/10 blend, so no retired cardinal behavior score enters the result. MiniMax is median-imputed because the index has no MiniMax entry.
5. Report uncertainty and coverage
Behavior uncertainty is estimated with 400 refits that resample whole transitive governed-source dependency components. Each draw carries its behavior ranks through the 90/10 percentile blend while holding governance fixed. These intervals measure sensitivity to the observed source corpus; they are not predictive intervals for future benchmarks and do not include uncertainty in benchmark design, company grades, governance mappings, or normative weights.
A model appears in the ranking only if it has direct results from at least 3 evals spanning at least 2 behavior components. This threshold is applied equally to every model.
Limitations
- Evaluation coverage is uneven: popular frontier models are tested more often than many open and Chinese models.
- Cross-era overlap remains weak; partial rankings propagate information through shared panels but cannot identify an ordering between disconnected model clusters.
- Prompts, judges, wrappers, deployment settings, and model aliases differ across sources.
- Published benchmark uncertainty is not propagated through the full index.
- Component weights are subjective value judgments.
- Company governance ≠ model behavior, and we also don't have governance data for every company.
Downloads
The public bundle contains the ranking and the minimum tables needed to understand its composition. Source links remain on the Evals page.
| File | Contents |
|---|---|
| rankings.csv | Overall ranks, bootstrap intervals, coverage, companies, and release dates |
| components.csv | Component weights and coverage |
| evals.csv | Evaluation names, constructs, model coverage, and source links |
| sub-evals.csv | Sub-eval directions, weights, and component mappings |