← Models

Model profile

Mistral Medium 3.1

Mistral AIdeveloper
2025-08-12release date
#209 / 267overall rank
5eval lineages

Evidence summary

Mistral Medium 3.1 has an estimated overall rank of #209; its 90% source-sensitivity interval is #116–#246. Its behavior-only rank is #201; company governance moves the combined estimate to #209. Published evidence spans 5 evals and 6 of 7 behavior components. Its strongest relative result is UAVBench safety-critical decision recognition (ethical_safety_critical_accuracy, #6 of 27); its weakest is SpeciEval (sea_animal_4ns, #100 of 102).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#163 / 3110.813↓ lowerSource ↗official
Alignment Leaderboardcorrigibility#17 / 244.117↑ higherSource ↗official
Alignment Leaderboardhonesty#20 / 243.212↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#22 / 242.665↑ higherSource ↗official
Alignment Leaderboardrobustness#19 / 243.147↑ higherSource ↗official
Alignment Leaderboardsafety#23 / 242.962↑ higherSource ↗official
Alignment Leaderboardscheming#21 / 243.319↑ higherSource ↗official
PHAREbias_resistance_diagnostic#45 / 660.3966↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#46 / 700.6963↑ higherSource ↗official
PHAREharm_resistance_diagnostic#40 / 700.9232↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#63 / 670.3083↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#33 / 1026.88↑ higherSource ↗official
SpeciEvalland_animal_4ns#89 / 1024.97↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#100 / 1025.58↓ lowerSource ↗official
SpeciEvalspeciesism#56 / 1022.09↓ lowerSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#6 / 270.75↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20
Government42.6
Diplomacy66.8
Economy43.8
Society61.1