← Models

Model profile

MiniMax M2

MiniMaxdeveloper
2025-10-27release date
#130 / 346Safety rank
#508 / 662Freedom rank

Evidence summary

Safety. MiniMax M2 has an estimated Safety rank of #130; its 90% source-sensitivity interval is #75–#285. Its behavior-only rank is #127; company governance moves the combined estimate to #130. Published Safety evidence spans 7 eval lineages and 6 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is SpeciEval (sea_animal_4ns, #122 of 131).

Freedom. MiniMax M2 has an estimated Freedom rank of #508; its 90% source-sensitivity interval is #362–#558. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — FRT-SciKnowEval-BiologicalHarmfulQA (safety_score, #11 of 45); its weakest is Concordia — Fortress-Privacy/Scams (safety_score, #48 of 54).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#295 / 358↓0.909Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#95 / 111↑1405.0Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#32 / 104↓18.05Source ↗official
Concordia — Agentic-Misalignmentsafety_score#23 / 54↑82.5Source ↗official
Concordia — AirBench-Deceptionsafety_score#19 / 63↑90.74Source ↗official
Concordia — AirBench-Manipulationsafety_score#12 / 56↑98Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#18 / 63↑98.2Source ↗official
Concordia — APEsafety_score#13 / 55↑58.36Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#51 / 63↑73.31Source ↗official
Concordia — DarkBenchsafety_score#10 / 55↑66.98Source ↗official
Concordia — Fortress-Biologicalsafety_score#12 / 54↑81.11Source ↗official
Concordia — Fortress-Chemicalsafety_score#14 / 54↑79.07Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#7 / 54↑79.81Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#11 / 45↑50Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#14 / 45↑38.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#33 / 45↑1.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#12 / 45↑65Source ↗official
Concordia — MASKsafety_score#14 / 62↑80.51Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#32 / 63↑82.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#18 / 63↑32.82Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53↑100Source ↗official
Concordia — SOSBench-Biosafety_score#27 / 63↑87.6Source ↗official
Concordia — SOSBench-Chemsafety_score#33 / 63↑86.8Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#18 / 270↑51.94Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#156 / 270↑86Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#76 / 270↑89.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#36 / 268↑99.09Source ↗official
LiveSecBenchethics#12 / 43↑69.21Source ↗official
LiveSecBenchfactuality#13 / 43↑62.29Source ↗official
LiveSecBenchlegality#12 / 43↑67.96Source ↗official
LiveSecBenchprivacy#14 / 43↑64.5Source ↗official
LiveSecBenchpsychological_health#10 / 43↑69.47Source ↗official
SpeciEvalbelief_animal_sentience#64 / 131↑6.82Source ↗official
SpeciEvalland_animal_4ns#118 / 131↓4.97Source ↗official
SpeciEvalsea_animal_4ns#122 / 131↓5.2Source ↗official
SpeciEvalspeciesism#47 / 131↓1.81Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#73 / 104↑18.05Source ↗official
Concordia — AirBench-Deceptionsafety_score#45 / 63↓90.74Source ↗official
Concordia — AirBench-Manipulationsafety_score#41 / 56↓98Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#46 / 56↓81.9Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#46 / 63↓98.2Source ↗official
Concordia — Fortress-Biologicalsafety_score#43 / 54↓81.11Source ↗official
Concordia — Fortress-Chemicalsafety_score#41 / 54↓79.07Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#48 / 54↓79.81Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#35 / 45↓50Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#32 / 45↓38.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#11 / 45↓1.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#34 / 45↓65Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#31 / 63↓82.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#45 / 63↓32.82Source ↗official
Concordia — SOSBench-Biosafety_score#37 / 63↓87.6Source ↗official
Concordia — SOSBench-Chemsafety_score#31 / 63↓86.8Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#113 / 270↓86Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#195 / 270↓89.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#232 / 268↓99.09Source ↗official
LiveSecBenchethics#32 / 43↓69.21Source ↗official
LiveSecBenchlegality#32 / 43↓67.96Source ↗official
LiveSecBenchprivacy#30 / 43↓64.5Source ↗official
LiveSecBenchpsychological_health#34 / 43↓69.47Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#61 / 156↑3Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-16.9
Government45.2
Diplomacy64
Economy48.6
Society59.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility55
Warning-phrase rate35
Soft-censorship rate0
API-error rate0

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions4.01
CCP-narrative alignment — China topics4.52
CCP-narrative alignment — non-China controls2.47