← Models

Model profile

MiMo v2 Flash

Xiaomideveloper
2025-12-16release date
#240 / 346Safety rank
#251 / 662Freedom rank

Evidence summary

Safety. MiMo v2 Flash has an estimated Safety rank of #240; its 90% source-sensitivity interval is #131–#290. Its behavior-only rank is #243; company governance moves the combined estimate to #240. Published Safety evidence spans 7 eval lineages and 6 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is SM-Bench (overfit, #91 of 92).

Freedom. MiMo v2 Flash has an estimated Freedom rank of #251; its 90% source-sensitivity interval is #110–#443. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (adversarial, #9 of 92); its weakest is SM-Bench (overfit, #91 of 92).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#78 / 358↓0.4841Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#60 / 111↑1433.0Source ↗official
BullshitBench v2clear_pushback_rate#92 / 122↑0.145Source ↗official
Concordia — Agentic-Misalignmentsafety_score#18 / 54↑87.58Source ↗official
Concordia — AirBench-Deceptionsafety_score#45 / 63↑72.22Source ↗official
Concordia — AirBench-Manipulationsafety_score#37 / 56↑80.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#53 / 63↑73.87Source ↗official
Concordia — APEsafety_score#46 / 55↑3.347Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#56 / 63↑70.12Source ↗official
Concordia — DarkBenchsafety_score#38 / 55↑49.85Source ↗official
Concordia — Fortress-Biologicalsafety_score#49 / 54↑26.82Source ↗official
Concordia — Fortress-Chemicalsafety_score#40 / 54↑33.67Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#47 / 54↑38.52Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#31 / 45↑26.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#31 / 45↑17.17Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#33 / 45↑1.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#32 / 45↑39.33Source ↗official
Concordia — MASKsafety_score#32 / 62↑59.47Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#57 / 63↑47.3Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#42 / 63↑12.78Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53↑100Source ↗official
Concordia — SOSBench-Biosafety_score#47 / 63↑59.92Source ↗official
Concordia — SOSBench-Chemsafety_score#54 / 63↑60.4Source ↗official
LiveSecBenchethics#6 / 43↑80.99Source ↗official
LiveSecBenchfactuality#19 / 43↑52.77Source ↗official
LiveSecBenchlegality#18 / 43↑52.92Source ↗official
LiveSecBenchprivacy#22 / 43↑53.69Source ↗official
LiveSecBenchpsychological_health#26 / 43↑45.77Source ↗official
SM-Benchadversarial#83 / 92↑73.66Source ↗official
SM-Benchambiguous_interpretation#65 / 92↑81.85Source ↗official
SM-Benchanti_hallucination#80 / 92↑80.63Source ↗official
SM-Bencheq_boundaries#52 / 92↑63.2Source ↗official
SM-Benchoverfit#91 / 92↑12.84Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#19 / 24↑33.55Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#19 / 63↓72.22Source ↗official
Concordia — AirBench-Manipulationsafety_score#20 / 56↓80.67Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#13 / 56↓44.76Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#11 / 63↓73.87Source ↗official
Concordia — Fortress-Biologicalsafety_score#6 / 54↓26.82Source ↗official
Concordia — Fortress-Chemicalsafety_score#15 / 54↓33.67Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#8 / 54↓38.52Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#15 / 45↓26.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#15 / 45↓17.17Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#11 / 45↓1.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#14 / 45↓39.33Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#7 / 63↓47.3Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#22 / 63↓12.78Source ↗official
Concordia — SOSBench-Biosafety_score#17 / 63↓59.92Source ↗official
Concordia — SOSBench-Chemsafety_score#10 / 63↓60.4Source ↗official
LiveSecBenchethics#38 / 43↓80.99Source ↗official
LiveSecBenchlegality#26 / 43↓52.92Source ↗official
LiveSecBenchprivacy#22 / 43↓53.69Source ↗official
LiveSecBenchpsychological_health#18 / 43↓45.77Source ↗official
SM-Benchadversarial#9 / 92↓73.66Source ↗official
SM-Bencheq_boundaries#52 / 92↑63.2Source ↗official
SM-Benchoverfit#91 / 92↑12.84Source ↗official
SpeechMap model completioncomplete_pct#108 / 181↑50.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#54 / 156↑3.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#46 / 156↑4Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-29.9
Government45
Diplomacy66.2
Economy51.5
Society66.3

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions4.45
CCP-narrative alignment — China topics4.76
CCP-narrative alignment — non-China controls3.49