← Models

Model profile

Claude Sonnet 4.6

Anthropicdeveloper
2026-01-21release date
#14 / 267overall rank
25eval lineages
1discovery sources

Evidence summary

Claude Sonnet 4.6 has an estimated overall rank of #14; its 90% source-sensitivity interval is #4–#72. Its behavior-only rank is #17; company governance moves the combined estimate to #14. Published evidence spans 25 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (anti_hallucination, #1 of 73); its weakest is Claude Sonnet 4.6 User Wellbeing (child_multiturn_appropriate_rate, #4 of 4).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#41 / 3110.4614↓ lowerSource ↗official
AgentAbstainabstain#4 / 1766.4↑ higherSource ↗official
AgentAbstaincar#4 / 1765.4↑ higherSource ↗official
AgentAbstainpaired#3 / 1753.4↑ higherSource ↗official
ANIMAscore#17 / 190.5646↑ higherSource ↗official
BioSecBench-Refusalbalanced_refusal_score#2 / 100.4304↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#2 / 1050.9↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#12 / 4354.8↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#6 / 4831.5↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#10 / 4943.1↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#17 / 4584.9↓ lowerSource ↗official
CAIS Risk Indexmask#14 / 5111.7↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#17 / 4816.5↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#9 / 1053.041↓ lowerSource ↗official
Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rate#2 / 50.18↓ lowerSource ↗official
Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rate#2 / 50.23↓ lowerSource ↗official
Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rate#2 / 40.08↓ lowerSource ↗official
Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rate#4 / 495↑ higherSource ↗official
Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rate#1 / 499.96↑ higherSource ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rate#3 / 40.17↓ lowerSource ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rate#2 / 499.73↑ higherSource ↗official
Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rate#1 / 498↑ higherSource ↗official
DystopiaBenchbasaglia_score#2 / 5024.2↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#6 / 5027.87↓ lowerSource ↗official
DystopiaBenchhuxley_score#3 / 5017.6↓ lowerSource ↗official
DystopiaBenchlaguardia_score#3 / 5033↓ lowerSource ↗official
DystopiaBenchorwell_score#5 / 5024.57↓ lowerSource ↗official
DystopiaBenchpetrov_score#3 / 5024.93↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#101 / 26018.86↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#184 / 26084.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#7 / 26099.44↑ higherSource ↗official
LiveSecBenchethics#8 / 4378.63↑ higherSource ↗official
LiveSecBenchfactuality#1 / 4392.17↑ higherSource ↗official
LiveSecBenchlegality#2 / 4397.77↑ higherSource ↗official
LiveSecBenchprivacy#3 / 4389.26↑ higherSource ↗official
LiveSecBenchpsychological_health#8 / 4372.01↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#17 / 5084.9↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#3 / 314.2↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
MORUscore#11 / 1369.73↑ higherSource ↗official
PHAREbias_resistance_diagnostic#42 / 660.4128↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#8 / 700.8551↑ higherSource ↗official
PHAREharm_resistance_diagnostic#2 / 700.9993↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#13 / 670.6939↑ higherSource ↗official
RefusalBenchyouden_j#5 / 190.6766↑ higherSource ↗official
SM-Benchadversarial#45 / 7380↑ higherSource ↗official
SM-Benchambiguous_interpretation#19 / 7388.99↑ higherSource ↗official
SM-Benchanti_hallucination#1 / 73100↑ higherSource ↗official
SM-Bencheq_boundaries#73 / 7333.43↑ higherSource ↗official
SM-Benchoverfit#10 / 7391.53↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#53 / 1026.78↑ higherSource ↗official
SpeciEvalland_animal_4ns#63 / 1024.65↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#38 / 1024.67↓ lowerSource ↗official
SpeciEvalspeciesism#57 / 1022.1↓ lowerSource ↗official
TACbase_welfare_rate#9 / 6837.82↑ higherSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#21 / 2510.9↓ lowerSource ↗official
Vigil Mental Health Safetyoverall_score#1 / 2383↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-22.1
Government46.3
Diplomacy64.6
Economy49
Society58.9

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.8
Honesty-humility7.3
Extraversion5.4
Agreeableness6.1
Conscientiousness7.1

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression1.56
Traditional ↔ Secular0.513