← Models

Model profile

Qwen3.6 Plus

Alibabadeveloper
2026-04-02release date
#164 / 346Safety rank
#622 / 662Freedom rank

Evidence summary

Safety. Qwen3.6 Plus has an estimated Safety rank of #164; its 90% source-sensitivity interval is #100–#275. Its behavior-only rank is #160; company governance moves the combined estimate to #164. Published Safety evidence spans 11 eval lineages and 7 of 7 components. Its strongest relative result is BullshitBench v2 (clear_pushback_rate, #15 of 122); its weakest is Every Model Cheats — Cybench Cheat Propensity (standard_anti_cheat_cheat_propensity_pct, #21 of 22).

Freedom. Qwen3.6 Plus has an estimated Freedom rank of #622; its 90% source-sensitivity interval is #415–#656. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is UGI Leaderboard — base-model willingness (willingness_adherence_score, #70 of 156); its weakest is SpeechMap model completion (complete_pct, #157 of 181).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#48 / 358↓0.3464Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#9 / 31↓23.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#10 / 31↓55.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#12 / 31↓47.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#6 / 31↓26.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#11 / 31↓41.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#8 / 31↓40Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#9 / 31↓43.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#7 / 31↓33.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#9 / 31↓31.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#10 / 31↓46.73Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#10 / 31↓34Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#10 / 31↓52.2Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#85 / 111↑1418.0Source ↗official
BullshitBench v2clear_pushback_rate#15 / 122↑0.655Source ↗official
DystopiaBenchbasaglia_score#38 / 50↓69Source ↗official
DystopiaBenchbaudrillard_score#44 / 50↓73.63Source ↗official
DystopiaBenchhuxley_score#42 / 50↓77.27Source ↗official
DystopiaBenchlaguardia_score#33 / 50↓68.67Source ↗official
DystopiaBenchorwell_score#38 / 50↓73.9Source ↗official
DystopiaBenchpetrov_score#39 / 50↓78.1Source ↗official
Every Model Cheats — Cybench Cheat Propensitybaseline_cheat_propensity_pct#16 / 22↓43.5Source ↗official
Every Model Cheats — Cybench Cheat Propensitysevere_anti_cheat_cheat_propensity_pct#14 / 22↓8.7Source ↗official
Every Model Cheats — Cybench Cheat Propensitystandard_anti_cheat_cheat_propensity_pct#21 / 22↓39.1Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#23 / 69↑0.3Source ↗official
SM-Benchadversarial#21 / 92↑85.85Source ↗official
SM-Benchambiguous_interpretation#24 / 92↑89.29Source ↗official
SM-Benchanti_hallucination#31 / 92↑96.86Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#55 / 92↑69.4Source ↗official
SpeciEvalbelief_animal_sentience#23 / 131↑6.98Source ↗official
SpeciEvalland_animal_4ns#69 / 131↓4.55Source ↗official
SpeciEvalsea_animal_4ns#54 / 131↓4.67Source ↗official
SpeciEvalspeciesism#94 / 131↓2.27Source ↗official
The Dictatorship Evaloverall_resistance_rate#15 / 20↑38.83Source ↗official
ToolPrivacyBenchprivate_mt_poi#7 / 9↓27.46Source ↗official
ToolPrivacyBenchpublic_mt_poi#7 / 9↓19.25Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#23 / 31↑23.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#22 / 31↑55.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#17 / 31↑47.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#25 / 31↑26.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#21 / 31↑41.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#24 / 31↑40Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#23 / 31↑43.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#25 / 31↑33.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#23 / 31↑31.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#22 / 31↑46.73Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#22 / 31↑34Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#22 / 31↑52.2Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#45 / 69↓0.3Source ↗official
SM-Benchadversarial#70 / 92↓85.85Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#55 / 92↑69.4Source ↗official
SpeechMap model completioncomplete_pct#157 / 181↑34.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#70 / 156↑2.25Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#99 / 156↑2.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.2
Government47
Diplomacy65.4
Economy45.1
Society61.5