← Models

Model profile

Qwen3 30B A3B Instruct

Alibabadeveloper
2025-07-28release date
#122 / 267overall rank
4eval lineages

Evidence summary

Qwen3 30B A3B Instruct has an estimated overall rank of #122; its 90% source-sensitivity interval is #43–#230. Its behavior-only rank is #117; company governance moves the combined estimate to #122. Published evidence spans 4 evals and 6 of 7 behavior components. Its strongest relative result is SpeciEval (speciesism, #14 of 102); its weakest is PHARE (hallucination_resistance_diagnostic, #68 of 70).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#150 / 3110.8029↓ lowerSource ↗official
PacifAIstp_score#3 / 788.89↑ higherSource ↗official
PHAREbias_resistance_diagnostic#17 / 660.5357↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#68 / 700.5988↑ higherSource ↗official
PHAREharm_resistance_diagnostic#63 / 700.8176↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#48 / 670.4209↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#60 / 1026.72↑ higherSource ↗official
SpeciEvalland_animal_4ns#86 / 1024.88↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#79 / 1025.03↓ lowerSource ↗official
SpeciEvalspeciesism#14 / 1021.5↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-24.8
Government45.9
Diplomacy68.2
Economy43.6
Society63.3