← Models

Model profile

Qwen3.5 397B A17B

Alibabadeveloper
2026-02-16release date
#53 / 346Safety rank
#526 / 662Freedom rank

Evidence summary

Safety. Qwen3.5 397B A17B has an estimated Safety rank of #53; its 90% source-sensitivity interval is #51–#116. Its behavior-only rank is #32; company governance moves the combined estimate to #53. Published Safety evidence spans 13 eval lineages and 7 of 7 components. Its strongest relative result is BioTIER (permit_compliance_pct, #1 of 52); its weakest is DelusionEval (discourages_harm_prevalence_pct, #15 of 16).

Freedom. Qwen3.5 397B A17B has an estimated Freedom rank of #526; its 90% source-sensitivity interval is #308–#594. Published Freedom evidence spans 9 eval lineages and 1 of 1 components. Its strongest relative result is BioTIER (permit_compliance_pct, #1 of 52); its weakest is LiveSecBench (ethics, #42 of 43).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#211 / 358↓0.8269Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#52 / 111↑1439.0Source ↗official
BioTIERpermit_compliance_pct#1 / 52↑100Source ↗official
BioTIERrefuse_compliance_pct#32 / 52↑29.5Source ↗official
BullshitBench v2clear_pushback_rate#9 / 122↑0.735Source ↗official
DelusionEvaldelusional_prevalence_pct#7 / 16↓28.65Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#15 / 16↑12.55Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#5 / 16↓1.15Source ↗official
DelusionEvalrelationship_prevalence_pct#6 / 16↓21.1Source ↗official
DelusionEvalsycophancy_prevalence_pct#4 / 16↓19.9Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#21 / 270↑50.39Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#216 / 270↑75.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#41 / 270↑97.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#95 / 268↑97.45Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#23 / 69↑0.3Source ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#6 / 10↑95.4Source ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#6 / 10↑77.3Source ↗official
Inkling-Small model card — StrongREJECTsafety_rate#1 / 10↑99.4Source ↗official
LiveSecBenchethics#2 / 43↑89.63Source ↗official
LiveSecBenchfactuality#8 / 43↑72.1Source ↗official
LiveSecBenchlegality#16 / 43↑62.2Source ↗official
LiveSecBenchprivacy#2 / 43↑94.2Source ↗official
LiveSecBenchpsychological_health#3 / 43↑89.45Source ↗official
SABERoverall_safety_rate#5 / 13↑26.58Source ↗official
SABERscenario_a_safety_rate#5 / 13↑30.62Source ↗official
SABERscenario_b_safety_rate#4 / 13↑36Source ↗official
SABERscenario_c_safety_rate#7 / 13↑15.02Source ↗official
SM-Benchadversarial#55 / 92↑80.49Source ↗official
SM-Benchambiguous_interpretation#71 / 92↑79.76Source ↗official
SM-Benchanti_hallucination#19 / 92↑98.43Source ↗official
SM-Bencheq_boundaries#68 / 92↑58.43Source ↗official
SM-Benchoverfit#68 / 92↑56.83Source ↗official
TACbase_welfare_rate#42 / 92↑28.21Source ↗self run
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#9 / 24↑40.53Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
BioTIERpermit_compliance_pct#1 / 52↑100Source ↗official
BioTIERrefuse_compliance_pct#21 / 52↓29.5Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#2 / 16↓12.55Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#12 / 16↑1.15Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#55 / 270↓75.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#229 / 270↓97.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#170 / 268↓97.45Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#45 / 69↓0.3Source ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#6 / 10↑95.4Source ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#5 / 10↓77.3Source ↗official
Inkling-Small model card — StrongREJECTsafety_rate#9 / 10↓99.4Source ↗official
LiveSecBenchethics#42 / 43↓89.63Source ↗official
LiveSecBenchlegality#28 / 43↓62.2Source ↗official
LiveSecBenchprivacy#42 / 43↓94.2Source ↗official
LiveSecBenchpsychological_health#41 / 43↓89.45Source ↗official
SM-Benchadversarial#35 / 92↓80.49Source ↗official
SM-Bencheq_boundaries#68 / 92↑58.43Source ↗official
SM-Benchoverfit#68 / 92↑56.83Source ↗official
SpeechMap model completioncomplete_pct#149 / 181↑37.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 156↑1.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.9
Government48
Diplomacy65.2
Economy44
Society57.9

CCPBench political narrative alignment

DimensionValueDistribution
CCP-narrative alignment — all questions3.21
CCP-narrative alignment — China topics3.87
CCP-narrative alignment — non-China controls1.22

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.5
Honesty-humility5.1
Extraversion6.3
Agreeableness6.7
Conscientiousness7.8

Agent-ValueBench Schwartz Basic Values (PVQ40)