← Models

Model profile

Qwen1.5 72B Chat

Alibabadeveloper
2024-02-04release date
#139 / 267overall rank
12eval lineages

Evidence summary

Qwen1.5 72B Chat has an estimated overall rank of #139; its 90% source-sensitivity interval is #58–#188. Its behavior-only rank is #133; company governance moves the combined estimate to #139. Published evidence spans 12 evals and 6 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #2 of 33); its weakest is CRiskEval (situational_awareness, #15 of 17).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#66 / 800.486↑ higherSource ↗official
ChineseSafescore#5 / 2263.67↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#2 / 140.22↓ lowerSource ↗official
ChiSafetyBenchmcq_score#1 / 1291.13↑ higherSource ↗official
CRiskEvaldeception_willingness#6 / 1720.12↓ lowerSource ↗official
CRiskEvaldesire_for_resource#7 / 1731.71↓ lowerSource ↗official
CRiskEvalharmful_goal#4 / 1733.39↓ lowerSource ↗official
CRiskEvalimprovement_intent#7 / 1748.6↓ lowerSource ↗official
CRiskEvalmalicious_coordination#6 / 178.07↓ lowerSource ↗official
CRiskEvalself_preservation#8 / 1736.86↓ lowerSource ↗official
CRiskEvalsituational_awareness#15 / 1768.75↓ lowerSource ↗official
HELM Safetyanthropic_red_team#37 / 800.99↑ higherSource ↗official
HELM Safetybbq#64 / 800.846↑ higherSource ↗official
HELM Safetyharmbench#56 / 800.648↑ higherSource ↗official
HELM Safetysimple_safety_tests#32 / 800.99↑ higherSource ↗official
HELM Safetyxstest#42 / 800.957↑ higherSource ↗official
OR-Benchover_refusal_rate#12 / 2546.9↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#16 / 255.6↓ lowerSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#13 / 3320.47↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#13 / 3317.92↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#13 / 3317.05↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#14 / 3318.42↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#16 / 3314.19↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#15 / 3314.29↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#14 / 3396.1↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#10 / 3396.89↑ higherSource ↗official
SALAD-Benchbase_malicious_use#15 / 3395.2↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#17 / 3393.65↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#16 / 3390.06↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#12 / 3391.89↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#2 / 3383.89↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#2 / 3380.28↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#2 / 3384.81↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#2 / 3380↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#2 / 3379.27↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#2 / 3378.89↑ higherSource ↗official
SORRY-Benchavg#34 / 510.36↓ lowerSource ↗official