← Models

Model profile

Qwen1.5 0.5B Chat

Alibabadeveloper
2024-02-04release date
#298 / 346Safety rank
#227 / 662Freedom rank

Evidence summary

Safety. Qwen1.5 0.5B Chat has an estimated Safety rank of #298; its 90% source-sensitivity interval is #77–#325. Its behavior-only rank is #290; company governance moves the combined estimate to #298. Published Safety evidence spans 3 eval lineages and 4 of 7 components. Its strongest relative result is Open LLM Safety Index (jailbreakbench_safety_rate, #1 of 21); its weakest is SALAD-Bench (mcq_representation_toxicity, #31 of 33).

Freedom. Qwen1.5 0.5B Chat has an estimated Freedom rank of #227; its 90% source-sensitivity interval is #104–#594. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #6 of 33); its weakest is Open LLM Safety Index (jailbreakbench_safety_rate, #20 of 21).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Open LLM Safety Indexjailbreakbench_safety_rate#1 / 21↑0.8Source ↗official
Open LLM Safety Indexstrongreject_safety_rate#1 / 21↑1Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#20 / 33↑11.42Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#25 / 33↑5.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#21 / 33↑8.4Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#25 / 33↑6.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#25 / 33↑6.79Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#18 / 33↑12.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#27 / 33↑83.23Source ↗official
SALAD-Benchbase_information_safety_harms#28 / 33↑84.9Source ↗official
SALAD-Benchbase_malicious_use#27 / 33↑80.92Source ↗official
SALAD-Benchbase_misinformation_harms#25 / 33↑88.08Source ↗official
SALAD-Benchbase_representation_toxicity#28 / 33↑75.3Source ↗official
SALAD-Benchbase_socioeconomic_harms#25 / 33↑81.2Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#30 / 33↑0Source ↗official
SALAD-Benchmcq_information_safety_harms#30 / 33↑0Source ↗official
SALAD-Benchmcq_malicious_use#31 / 33↑0Source ↗official
SALAD-Benchmcq_misinformation_harms#30 / 33↑0Source ↗official
SALAD-Benchmcq_representation_toxicity#31 / 33↑0Source ↗official
SALAD-Benchmcq_socioeconomic_harms#30 / 33↑0Source ↗official
SORRY-Benchavg#41 / 51↓0.43Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Open LLM Safety Indexjailbreakbench_safety_rate#20 / 21↓0.8Source ↗official
Open LLM Safety Indexstrongreject_safety_rate#19 / 21↓1Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#14 / 33↓11.42Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#7 / 33↓5.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#12 / 33↓8.4Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#9 / 33↓6.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#9 / 33↓6.79Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#15 / 33↓12.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#7 / 33↓83.23Source ↗official
SALAD-Benchbase_information_safety_harms#6 / 33↓84.9Source ↗official
SALAD-Benchbase_malicious_use#7 / 33↓80.92Source ↗official
SALAD-Benchbase_misinformation_harms#9 / 33↓88.08Source ↗official
SALAD-Benchbase_representation_toxicity#6 / 33↓75.3Source ↗official
SALAD-Benchbase_socioeconomic_harms#9 / 33↓81.2Source ↗official
SORRY-Benchavg#11 / 51↑0.43Source ↗official