← Models

Model profile

Qwen3.5 Plus

Alibabadeveloper
2026-02-15release date
#81 / 346Safety rank
#637 / 662Freedom rank

Evidence summary

Safety. Qwen3.5 Plus has an estimated Safety rank of #81; its 90% source-sensitivity interval is #59–#176. Its behavior-only rank is #68; company governance moves the combined estimate to #81. Published Safety evidence spans 7 eval lineages and 6 of 7 components. Its strongest relative result is Concordia — AirBench-SecurityRisks (safety_score, #1 of 63); its weakest is SM-Bench (overfit, #84 of 92).

Freedom. Qwen3.5 Plus has an estimated Freedom rank of #637; its 90% source-sensitivity interval is #476–#652. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (adversarial, #25 of 92); its weakest is Concordia — FRT-SOSBench-Chem (safety_score, #43 of 45).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#11 / 31↓27.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#9 / 31↓51.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#9 / 31↓39.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#8 / 31↓28.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#8 / 31↓33.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#10 / 31↓47.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#8 / 31↓42.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#8 / 31↓35.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#10 / 31↓34.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#9 / 31↓42.48Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#9 / 31↓29.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#9 / 31↓41.3Source ↗official
Concordia — Agentic-Misalignmentsafety_score#16 / 54↑87.83Source ↗official
Concordia — AirBench-Deceptionsafety_score#13 / 63↑95.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#12 / 56↑98Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#1 / 63↑100Source ↗official
Concordia — APEsafety_score#11 / 55↑61.14Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#8 / 63↑96.41Source ↗official
Concordia — DarkBenchsafety_score#17 / 55↑60.15Source ↗official
Concordia — Fortress-Biologicalsafety_score#18 / 54↑76.75Source ↗official
Concordia — Fortress-Chemicalsafety_score#13 / 54↑79.12Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#12 / 54↑76.1Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#7 / 45↑55.33Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#5 / 45↑50.33Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#18 / 45↑3.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#3 / 45↑85.67Source ↗official
Concordia — MASKsafety_score#17 / 62↑79.27Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#14 / 63↑98.32Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#7 / 63↑47.14Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53↑100Source ↗official
Concordia — SOSBench-Biosafety_score#6 / 63↑98.8Source ↗official
Concordia — SOSBench-Chemsafety_score#4 / 63↑98.6Source ↗official
LiveSecBenchethics#5 / 43↑84.35Source ↗official
LiveSecBenchfactuality#2 / 43↑92.1Source ↗official
LiveSecBenchlegality#7 / 43↑84.85Source ↗official
LiveSecBenchprivacy#6 / 43↑79.84Source ↗official
LiveSecBenchpsychological_health#5 / 43↑80.56Source ↗official
MT-JailBench CrescendoXsafety_score#4 / 21↑47.8Source ↗official
SM-Benchadversarial#64 / 92↑79.51Source ↗official
SM-Benchambiguous_interpretation#73 / 92↑78.57Source ↗official
SM-Benchanti_hallucination#41 / 92↑95.29Source ↗official
SM-Bencheq_boundaries#58 / 92↑62.08Source ↗official
SM-Benchoverfit#84 / 92↑28.96Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#13 / 23↑85.94Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#61 / 94↑89.3Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#21 / 31↑27.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#23 / 31↑51.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#23 / 31↑39.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#24 / 31↑28.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#24 / 31↑33.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#22 / 31↑47.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#24 / 31↑42.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#24 / 31↑35.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#22 / 31↑34.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#23 / 31↑42.48Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#23 / 31↑29.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#23 / 31↑41.3Source ↗official
Concordia — AirBench-Deceptionsafety_score#50 / 63↓95.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#41 / 56↓98Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#47 / 56↓82.38Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#58 / 63↓100Source ↗official
Concordia — Fortress-Biologicalsafety_score#37 / 54↓76.75Source ↗official
Concordia — Fortress-Chemicalsafety_score#42 / 54↓79.12Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#43 / 54↓76.1Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#39 / 45↓55.33Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#41 / 45↓50.33Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#27 / 45↓3.667Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#43 / 45↓85.67Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#49 / 63↓98.32Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#57 / 63↓47.14Source ↗official
Concordia — SOSBench-Biosafety_score#56 / 63↓98.8Source ↗official
Concordia — SOSBench-Chemsafety_score#60 / 63↓98.6Source ↗official
LiveSecBenchethics#39 / 43↓84.35Source ↗official
LiveSecBenchlegality#37 / 43↓84.85Source ↗official
LiveSecBenchprivacy#38 / 43↓79.84Source ↗official
LiveSecBenchpsychological_health#39 / 43↓80.56Source ↗official
MT-JailBench CrescendoXsafety_score#18 / 21↓47.8Source ↗official
SM-Benchadversarial#25 / 92↓79.51Source ↗official
SM-Bencheq_boundaries#58 / 92↑62.08Source ↗official
SM-Benchoverfit#84 / 92↑28.96Source ↗official