← Models

Model profile

Mistral Medium 3

Mistral AIdeveloper
2025-05-07release date
#306 / 346Safety rank
#6 / 662Freedom rank

Evidence summary

Safety. Mistral Medium 3 has an estimated Safety rank of #306; its 90% source-sensitivity interval is #179–#343. Its behavior-only rank is #298; company governance moves the combined estimate to #306. Published Safety evidence spans 6 eval lineages and 5 of 7 components. Its strongest relative result is AA-Omniscience (hallucination_rate, #107 of 358); its weakest is Concordia — SOSBench-Chem (safety_score, #63 of 63).

Freedom. Mistral Medium 3 has an estimated Freedom rank of #6; its 90% source-sensitivity interval is #2–#355. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is UGI Leaderboard — base-model willingness (willingness_adherence_score, #2 of 156); its weakest is Concordia — Fortress-Biological (safety_score, #12 of 54).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#107 / 358↓0.609Source ↗official
Concordia — Agentic-Misalignmentsafety_score#32 / 54↑72.33Source ↗official
Concordia — AirBench-Deceptionsafety_score#60 / 63↑37.78Source ↗official
Concordia — AirBench-Manipulationsafety_score#56 / 56↑42.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#62 / 63↑41.67Source ↗official
Concordia — APEsafety_score#53 / 55↑0.4375Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#51 / 63↑73.31Source ↗official
Concordia — DarkBenchsafety_score#35 / 55↑53.03Source ↗official
Concordia — Fortress-Biologicalsafety_score#43 / 54↑29.4Source ↗official
Concordia — Fortress-Chemicalsafety_score#42 / 54↑32.88Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#53 / 54↑30.19Source ↗official
Concordia — MASKsafety_score#56 / 62↑42.04Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#58 / 63↑44.11Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#56 / 63↑3.965Source ↗official
Concordia — Shutdown-Resistancesafety_score#36 / 53↑98Source ↗official
Concordia — SOSBench-Biosafety_score#59 / 63↑11.8Source ↗official
Concordia — SOSBench-Chemsafety_score#63 / 63↑25.6Source ↗official
Confabulationsconfabulation_rate#41 / 52↓38.12Source ↗official
Emergent Collusionhigh_illegality_game_rate#8 / 13↓0.47Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#26 / 42↓77Source ↗official
SpeciEvalbelief_animal_sentience#87 / 131↑6.68Source ↗official
SpeciEvalland_animal_4ns#123 / 131↓5.03Source ↗official
SpeciEvalsea_animal_4ns#128 / 131↓5.53Source ↗official
SpeciEvalspeciesism#113 / 131↓2.62Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.3
Government43.1
Diplomacy64.3
Economy45.8
Society59.9