← Models

Model profile

Qwen1.5 7B Chat

Alibabadeveloper
2024-02-04release date
#169 / 267overall rank
9eval lineages

Evidence summary

Qwen1.5 7B Chat has an estimated overall rank of #169; its 90% source-sensitivity interval is #86–#203. Its behavior-only rank is #159; company governance moves the combined estimate to #169. Published evidence spans 9 evals and 6 of 7 behavior components. Its strongest relative result is CRiskEval (malicious_coordination, #1 of 17); its weakest is SALAD-Bench (attack_enhanced_socioeconomic_harms, #27 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#6 / 2262.48↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#5 / 140.43↓ lowerSource ↗official
ChiSafetyBenchmcq_score#8 / 1279.39↑ higherSource ↗official
Contextual MoralChoicehuman_agreement#12 / 220.4↑ higherSource ↗official
CRiskEvaldeception_willingness#9 / 1722.6↓ lowerSource ↗official
CRiskEvaldesire_for_resource#8 / 1732.14↓ lowerSource ↗official
CRiskEvalharmful_goal#9 / 1737.92↓ lowerSource ↗official
CRiskEvalimprovement_intent#6 / 1747.47↓ lowerSource ↗official
CRiskEvalmalicious_coordination#1 / 175.68↓ lowerSource ↗official
CRiskEvalself_preservation#7 / 1736.01↓ lowerSource ↗official
CRiskEvalsituational_awareness#13 / 1767.2↓ lowerSource ↗official
HyperCLOVA X Toxic Continuation Panelskold_toxic_count#2 / 70.0036↓ lowerSource ↗official
HyperCLOVA X Toxic Continuation Panelskold_toxicity#2 / 70.1061↓ lowerSource ↗official
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_count#3 / 70.006↓ lowerSource ↗official
HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicity#3 / 70.055↓ lowerSource ↗official
JailBenchjailbreak_success_rate#11 / 1471.6↓ lowerSource ↗official
OR-Benchover_refusal_rate#10 / 2539.2↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#18 / 2515↓ lowerSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#22 / 339.91↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#21 / 337.17↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#23 / 338.32↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#19 / 339.38↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#23 / 337.39↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#27 / 337.36↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#16 / 3395.17↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#14 / 3395.67↑ higherSource ↗official
SALAD-Benchbase_malicious_use#16 / 3394.12↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#14 / 3394.04↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#15 / 3390.23↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#13 / 3391.19↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#24 / 3319.17↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#26 / 3311.94↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#24 / 3316.35↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#24 / 3318.81↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#25 / 3315.62↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#24 / 3317.78↑ higherSource ↗official
SORRY-Benchavg#39 / 510.4↓ lowerSource ↗official