← Models

Model profile

Qwen1.5 32B Chat

Alibabadeveloper
2024-04-02release date
#180 / 346Safety rank
#425 / 662Freedom rank

Evidence summary

Safety. Qwen1.5 32B Chat has an estimated Safety rank of #180; its 90% source-sensitivity interval is #101–#234. Its behavior-only rank is #173; company governance moves the combined estimate to #180. Published Safety evidence spans 4 eval lineages and 6 of 7 components. Its strongest relative result is ChiSafetyBench (harmful_response_rate, #2 of 14); its weakest is CRiskEval (situational_awareness, #10 of 17).

Freedom. Qwen1.5 32B Chat has an estimated Freedom rank of #425; its 90% source-sensitivity interval is #250–#592. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is OR-Bench (over_refusal_rate, #13 of 25); its weakest is ChiSafetyBench (refusal_rr1, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#2 / 14↓0.22Source ↗official
ChiSafetyBenchmcq_score#4 / 12↑87.3Source ↗official
CRiskEvaldeception_willingness#5 / 17↓19.92Source ↗official
CRiskEvaldesire_for_resource#6 / 17↓28.56Source ↗official
CRiskEvalharmful_goal#7 / 17↓35.65Source ↗official
CRiskEvalimprovement_intent#8 / 17↓48.96Source ↗official
CRiskEvalmalicious_coordination#9 / 17↓10.48Source ↗official
CRiskEvalself_preservation#6 / 17↓29.81Source ↗official
CRiskEvalsituational_awareness#10 / 17↓65.38Source ↗official
OR-Benchover_refusal_rate#13 / 25↓50.8Source ↗official
OR-Benchtoxic_acceptance_rate#13 / 25↓4.4Source ↗official
SORRY-Benchavg#25 / 51↓0.28Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#11 / 14↑0.22Source ↗official
ChiSafetyBenchrefusal_rr1#14 / 14↓77.71Source ↗official
ChiSafetyBenchrefusal_rr2#14 / 14↓77.27Source ↗official
OR-Benchover_refusal_rate#13 / 25↓50.8Source ↗official
OR-Benchtoxic_acceptance_rate#13 / 25↑4.4Source ↗official
SORRY-Benchavg#27 / 51↑0.28Source ↗official