← Models

Model profile

Qwen 2.5 Max

Alibabadeveloper
2025-01-25release date
#221 / 267overall rank
4eval lineages

Evidence summary

Qwen 2.5 Max has an estimated overall rank of #221; its 90% source-sensitivity interval is #147–#248. Its behavior-only rank is #217; company governance moves the combined estimate to #221. Published evidence spans 4 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #52 of 260); its weakest is SM-Bench (eq_boundaries, #69 of 73).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Confabulationsconfabulation_rate#38 / 5231.19↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#123 / 26016.28↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#52 / 26092.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#142 / 26067.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#76 / 25898.09↑ higherSource ↗official
PHAREbias_resistance_diagnostic#40 / 660.4295↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#45 / 700.698↑ higherSource ↗official
PHAREharm_resistance_diagnostic#50 / 700.8989↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#39 / 670.4626↑ higherSource ↗official
SM-Benchadversarial#69 / 7370.73↑ higherSource ↗official
SM-Benchambiguous_interpretation#59 / 7377.38↑ higherSource ↗official
SM-Benchanti_hallucination#46 / 7388.48↑ higherSource ↗official
SM-Bencheq_boundaries#69 / 7347.47↑ higherSource ↗official
SM-Benchoverfit#42 / 7366.67↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism79.8
Self-direction56.1
Care / Harm28.6
Fairness / Cheating27.2
Ethical89.7