Model profile
Mistral Large 2
Evidence summary
Mistral Large 2 has an estimated overall rank of #262; its 90% source-sensitivity interval is #228–#264. Its behavior-only rank is #255; company governance moves the combined estimate to #262. Published evidence spans 11 evals and 5 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #92 of 311); its weakest is AgentHarm (harm_score, #12 of 12).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗mistral-large-2
- Hugging Face ↗mistralai/Mistral-Large-Instruct-2407
- OpenRouter ↗mistralai/mistral-large-2407
- Official model page ↗Exact model document · Reviewed official Hugging Face owner · official repository
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #92 / 311 | 0.6784 | ↓ lower | Source ↗official | |
| AgentHarmharm_score | #12 / 12 | 82.2 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #28 / 32 | 35.9 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #29 / 32 | 39.2 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chathate | #29 / 32 | 24.7 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #28 / 32 | 49.4 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #27 / 32 | 25.4 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #29 / 32 | 39.4 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #29 / 32 | 31.7 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #28 / 32 | 34.4 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #28 / 32 | 48.6 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #29 / 32 | 33.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #27 / 32 | 27.3 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #28 / 32 | 45.7 | ↓ lower | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #75 / 80 | 0.353 | ↑ higher | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #88 / 105 | 69.79 | ↓ lower | Source ↗official | |
| Confabulationsconfabulation_rate | #39 / 52 | 32.18 | ↓ lower | Source ↗official | |
| LiveSecBenchethics | #33 / 43 | 27.9 | ↑ higher | Source ↗official | |
| LiveSecBenchfactuality | #37 / 43 | 26.5 | ↑ higher | Source ↗official | |
| LiveSecBenchlegality | #24 / 43 | 42.83 | ↑ higher | Source ↗official | |
| LiveSecBenchprivacy | #34 / 43 | 26.82 | ↑ higher | Source ↗official | |
| LiveSecBenchpsychological_health | #37 / 43 | 24.54 | ↑ higher | Source ↗official | |
| MASKlying_probability_pct | #46 / 53 | 52.47 | ↓ lower | Source ↗official | |
| PHAREbias_resistance_diagnostic | #44 / 66 | 0.397 | ↑ higher | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #33 / 70 | 0.7562 | ↑ higher | Source ↗official | |
| PHAREharm_resistance_diagnostic | #52 / 70 | 0.8938 | ↑ higher | Source ↗official | |
| PHAREjailbreak_resistance_diagnostic | #64 / 67 | 0.3058 | ↑ higher | Source ↗official | |
| Shelleducation_jsr | #11 / 14 | 0.79 | ↓ lower | Source ↗official | |
| Shellfinance_jsr | #13 / 14 | 0.912 | ↓ lower | Source ↗official | |
| Shellmanagement_jsr | #13 / 14 | 0.92 | ↓ lower | Source ↗official | |
| SORRY-Benchavg | #45 / 51 | 0.6 | ↓ lower | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
