← Models

Model profile

Qwen1.5 14B Chat

Alibabadeveloper
2024-02-04release date
#194 / 346Safety rank
#214 / 662Freedom rank

Evidence summary

Safety. Qwen1.5 14B Chat has an estimated Safety rank of #194; its 90% source-sensitivity interval is #101–#255. Its behavior-only rank is #184; company governance moves the combined estimate to #194. Published Safety evidence spans 6 eval lineages and 5 of 7 components. Its strongest relative result is CRiskEval (malicious_coordination, #2 of 17); its weakest is SALAD-Bench (attack_enhanced_misinformation_harms, #31 of 33).

Freedom. Qwen1.5 14B Chat has an estimated Freedom rank of #214; its 90% source-sensitivity interval is #121–#316. Published Freedom evidence spans 4 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #21 of 268); its weakest is SALAD-Bench (base_socioeconomic_harms, #26 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChineseSafescore#7 / 22↑61.29Source ↗official
ChiSafetyBenchharmful_response_rate#2 / 14↓0.22Source ↗official
ChiSafetyBenchmcq_score#3 / 12↑88.39Source ↗official
CRiskEvaldeception_willingness#11 / 17↓23.53Source ↗official
CRiskEvaldesire_for_resource#5 / 17↓28.02Source ↗official
CRiskEvalharmful_goal#5 / 17↓34.36Source ↗official
CRiskEvalimprovement_intent#5 / 17↓46.04Source ↗official
CRiskEvalmalicious_coordination#2 / 17↓6.52Source ↗official
CRiskEvalself_preservation#5 / 17↓29.41Source ↗official
CRiskEvalsituational_awareness#12 / 17↓66.62Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#106 / 270↑19.12Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#150 / 270↑86.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#180 / 270↑57.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#248 / 268↑80.95Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#26 / 33↑7.11Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#30 / 33↑3.91Source ↗official
SALAD-Benchattack_enhanced_malicious_use#26 / 33↑6.53Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#31 / 33↑4.28Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#26 / 33↑6.19Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#22 / 33↑9.09Source ↗official
SALAD-Benchbase_human_autonomy_integrity#14 / 33↑96.1Source ↗official
SALAD-Benchbase_information_safety_harms#8 / 33↑97.77Source ↗official
SALAD-Benchbase_malicious_use#10 / 33↑96.87Source ↗official
SALAD-Benchbase_misinformation_harms#10 / 33↑95.72Source ↗official
SALAD-Benchbase_representation_toxicity#11 / 33↑92.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#8 / 33↑93.65Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#6 / 33↑62.22Source ↗official
SALAD-Benchmcq_information_safety_harms#9 / 33↑49.72Source ↗official
SALAD-Benchmcq_malicious_use#7 / 33↑55.77Source ↗official
SALAD-Benchmcq_misinformation_harms#8 / 33↑54.52Source ↗official
SALAD-Benchmcq_representation_toxicity#7 / 33↑57.71Source ↗official
SALAD-Benchmcq_socioeconomic_harms#7 / 33↑52.78Source ↗official
SORRY-Benchavg#30 / 51↓0.34Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#11 / 14↑0.22Source ↗official
ChiSafetyBenchrefusal_rr1#8 / 14↓73.16Source ↗official
ChiSafetyBenchrefusal_rr2#8 / 14↓73.16Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#120 / 270↓86.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#89 / 270↓57.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#21 / 268↓80.95Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#8 / 33↓7.11Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#3 / 33↓3.91Source ↗official
SALAD-Benchattack_enhanced_malicious_use#8 / 33↓6.53Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#3 / 33↓4.28Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#8 / 33↓6.19Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#11 / 33↓9.09Source ↗official
SALAD-Benchbase_human_autonomy_integrity#19 / 33↓96.1Source ↗official
SALAD-Benchbase_information_safety_harms#26 / 33↓97.77Source ↗official
SALAD-Benchbase_malicious_use#24 / 33↓96.87Source ↗official
SALAD-Benchbase_misinformation_harms#24 / 33↓95.72Source ↗official
SALAD-Benchbase_representation_toxicity#22 / 33↓92.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#26 / 33↓93.65Source ↗official
SORRY-Benchavg#22 / 51↑0.34Source ↗official