← Models

Model profile

Qwen1.5 32B Chat

Alibabadeveloper
2024-04-02release date
#129 / 267overall rank
4eval lineages

Evidence summary

Qwen1.5 32B Chat has an estimated overall rank of #129; its 90% source-sensitivity interval is #58–#187. Its behavior-only rank is #124; company governance moves the combined estimate to #129. Published evidence spans 4 evals and 6 of 7 behavior components. Its strongest relative result is ChiSafetyBench (harmful_response_rate, #2 of 14); its weakest is CRiskEval (situational_awareness, #10 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChiSafetyBenchharmful_response_rate#2 / 140.22↓ lowerSource ↗official
ChiSafetyBenchmcq_score#4 / 1287.3↑ higherSource ↗official
CRiskEvaldeception_willingness#5 / 1719.92↓ lowerSource ↗official
CRiskEvaldesire_for_resource#6 / 1728.56↓ lowerSource ↗official
CRiskEvalharmful_goal#7 / 1735.65↓ lowerSource ↗official
CRiskEvalimprovement_intent#8 / 1748.96↓ lowerSource ↗official
CRiskEvalmalicious_coordination#9 / 1710.48↓ lowerSource ↗official
CRiskEvalself_preservation#6 / 1729.81↓ lowerSource ↗official
CRiskEvalsituational_awareness#10 / 1765.38↓ lowerSource ↗official
OR-Benchover_refusal_rate#13 / 2550.8↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#13 / 254.4↓ lowerSource ↗official
SORRY-Benchavg#25 / 510.28↓ lowerSource ↗official