← Models

Model profile

Magistral Medium

Mistral AIdeveloper
2025-06-10release date
#244 / 346Safety rank
#5 / 662Freedom rank

Evidence summary

Safety. Magistral Medium has an estimated Safety rank of #244; its 90% source-sensitivity interval is #145–#323. Its behavior-only rank is #230; company governance moves the combined estimate to #244. Published Safety evidence spans 8 eval lineages and 6 of 7 components. Its strongest relative result is FlagEval Safety and Values (a3_qualified_rate, #3 of 18); its weakest is Adversarial Poetry — AILuminate Baseline and Poetry ASR (baseline_asr, #24 of 24).

Freedom. Magistral Medium has an estimated Freedom rank of #5; its 90% source-sensitivity interval is #16–#106. Published Freedom evidence spans 4 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is PHARE (jailbreak_resistance_diagnostic, #27 of 67).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#105 / 358↓0.5966Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#24 / 24↓22.92Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#24 / 24↓77.19Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#141 / 270↑15.76Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#181 / 270↑83.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#252 / 270↑33.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#214 / 268↑91.82Source ↗official
FlagEval Safety and Valuesa1_qualified_rate#12 / 18↑79.01Source ↗official
FlagEval Safety and Valuesa2_qualified_rate#12 / 18↑78.73Source ↗official
FlagEval Safety and Valuesa3_qualified_rate#3 / 18↑89.28Source ↗official
FlagEval Safety and Valuesa4_qualified_rate#7 / 18↑89.15Source ↗official
FlagEval Safety and Valuesa5_qualified_rate#3 / 18↑74.27Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#36 / 54↑26.12Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
MASKlying_probability_pct#31 / 53↓43.5Source ↗official
PHAREbias_resistance_diagnostic#25 / 66↑0.5075Source ↗official
PHAREhallucination_resistance_diagnostic#49 / 70↑0.6819Source ↗official
PHAREharm_resistance_diagnostic#60 / 70↑0.8452Source ↗official
PHAREjailbreak_resistance_diagnostic#41 / 67↑0.4561Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#1 / 24↑22.92Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#1 / 24↑77.19Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#90 / 270↓83.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#18 / 270↓33.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#54 / 268↓91.82Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
PHAREharm_resistance_diagnostic#11 / 70↓0.8452Source ↗official
PHAREjailbreak_resistance_diagnostic#27 / 67↓0.4561Source ↗official