← Models

Model profile

Qwen3 14B

Alibabadeveloper
2025-04-27release date
#241 / 267overall rank
3eval lineages

Evidence summary

Qwen3 14B has an estimated overall rank of #241; its 90% source-sensitivity interval is #183–#258. Its behavior-only rank is #239; company governance moves the combined estimate to #241. Published evidence spans 3 evals and 6 of 7 behavior components. Its strongest relative result is Alignment Leaderboard (corrigibility, #13 of 24); its weakest is Alignment Leaderboard (robustness, #23 of 24).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Alignment Leaderboardcorrigibility#13 / 244.204↑ higherSource ↗official
Alignment Leaderboardhonesty#22 / 243.147↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#21 / 243.013↑ higherSource ↗official
Alignment Leaderboardrobustness#23 / 242.813↑ higherSource ↗official
Alignment Leaderboardsafety#19 / 243.558↑ higherSource ↗official
Alignment Leaderboardscheming#18 / 243.458↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#28 / 460.97↑ higherSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#18 / 2510.2↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20
Government49.1
Diplomacy64.9
Economy46.9
Society61.5