← Models

Model profile

Claude Opus 4

Anthropicdeveloper
2025-05-22release date
#6 / 267overall rank
19eval lineages

Evidence summary

Claude Opus 4 has an estimated overall rank of #6; its 90% source-sensitivity interval is #3–#56. Its behavior-only rank is #9; company governance moves the combined estimate to #6. Published evidence spans 19 evals and 7 of 7 behavior components. Its strongest relative result is Confabulations (confabulation_rate, #1 of 52); its weakest is Anthropic Claude Opus 4.1 System Card Addendum (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#11 / 800.857↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#15 / 1696↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#8 / 1657↓ lowerSource ↗official
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct#5 / 1065↓ lowerSource ↗official
Anthropic Claude 4 System Cardagentic_coding_safety#2 / 30.88↑ higherSource ↗official
Anthropic Claude 4 System Cardbenign_request_refusal#1 / 30.0007↓ lowerSource ↗official
Anthropic Claude 4 System Cardharmful_request_safety#3 / 30.9843↑ higherSource ↗official
Anthropic Claude 4 System Cardstrongreject_jailbreak_success#2 / 30.0714↓ lowerSource ↗official
Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracy#1 / 20.911↑ higherSource ↗official
Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal#1 / 20.0005↓ lowerSource ↗official
Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety#2 / 20.9727↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#43 / 1050.34↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#16 / 1054.736↓ lowerSource ↗official
Confabulationsconfabulation_rate#1 / 522.723↓ lowerSource ↗official
Emergent Collusionhigh_illegality_game_rate#6 / 130.36↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#27 / 26043.41↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#236 / 26075↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#17 / 26098.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#41 / 25899.18↑ higherSource ↗official
FORTRESSaverage_risk_score#22 / 4926.19↓ lowerSource ↗official
FORTRESSover_refusal_score#13 / 462.205↓ lowerSource ↗official
HELM Safetyanthropic_red_team#40 / 800.989↑ higherSource ↗official
HELM Safetybbq#4 / 800.982↑ higherSource ↗official
HELM Safetyharmbench#21 / 800.904↑ higherSource ↗official
HELM Safetysimple_safety_tests#26 / 800.9975↑ higherSource ↗official
HELM Safetyxstest#24 / 800.97↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#42 / 5424.9↑ higherSource ↗official
MASKlying_probability_pct#12 / 5315.93↓ lowerSource ↗official
Social Welfare Function Benchmarkfairness#5 / 190.547↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#71 / 1026.58↑ higherSource ↗official
SpeciEvalland_animal_4ns#37 / 1024.42↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#22 / 1024.53↓ lowerSource ↗official
SpeciEvalspeciesism#49 / 1021.98↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-21.7
Government44.4
Diplomacy67.6
Economy47.2
Society61.2

CAISI CCP narrative alignment

DimensionValueDistribution
CCP narrative alignment3.25