← Models

Model profile

Mistral Large 2

Mistral AIdeveloper
2024-07-24release date
#262 / 267overall rank
11eval lineages

Evidence summary

Mistral Large 2 has an estimated overall rank of #262; its 90% source-sensitivity interval is #228–#264. Its behavior-only rank is #255; company governance moves the combined estimate to #262. Published evidence spans 11 evals and 5 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #92 of 311); its weakest is AgentHarm (harm_score, #12 of 12).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#92 / 3110.6784↓ lowerSource ↗official
AgentHarmharm_score#12 / 1282.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#28 / 3235.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#29 / 3239.2↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#29 / 3224.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#28 / 3249.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#27 / 3225.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#29 / 3239.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#29 / 3231.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#28 / 3234.4↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#28 / 3248.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#29 / 3233.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#27 / 3227.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#28 / 3245.7↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#75 / 800.353↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#88 / 10569.79↓ lowerSource ↗official
Confabulationsconfabulation_rate#39 / 5232.18↓ lowerSource ↗official
LiveSecBenchethics#33 / 4327.9↑ higherSource ↗official
LiveSecBenchfactuality#37 / 4326.5↑ higherSource ↗official
LiveSecBenchlegality#24 / 4342.83↑ higherSource ↗official
LiveSecBenchprivacy#34 / 4326.82↑ higherSource ↗official
LiveSecBenchpsychological_health#37 / 4324.54↑ higherSource ↗official
MASKlying_probability_pct#46 / 5352.47↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#44 / 660.397↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#33 / 700.7562↑ higherSource ↗official
PHAREharm_resistance_diagnostic#52 / 700.8938↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#64 / 670.3058↑ higherSource ↗official
Shelleducation_jsr#11 / 140.79↓ lowerSource ↗official
Shellfinance_jsr#13 / 140.912↓ lowerSource ↗official
Shellmanagement_jsr#13 / 140.92↓ lowerSource ↗official
SORRY-Benchavg#45 / 510.6↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-21.6
Government44.8
Diplomacy67
Economy42.8
Society61.6

ValueCompass

DimensionValueDistribution
Universalism71.3
Self-direction51.9
Care / Harm34.6
Fairness / Cheating32.2
Ethical88.1