← Models

Model profile

Claude Opus 4

Anthropicdeveloper
2025-05-22release date
#42 / 346Safety rank
#532 / 662Freedom rank

Evidence summary

Safety. Claude Opus 4 has an estimated Safety rank of #42; its 90% source-sensitivity interval is #19–#124. Its behavior-only rank is #51; company governance moves the combined estimate to #42. Published Safety evidence spans 26 eval lineages and 7 of 7 components. Its strongest relative result is Confabulations (confabulation_rate, #1 of 52); its weakest is Anthropic Claude Opus 4.1 System Card Addendum (harmful_request_safety, #2 of 2).

Freedom. Claude Opus 4 has an estimated Freedom rank of #532; its 90% source-sensitivity interval is #369–#610. Published Freedom evidence spans 13 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #53 of 270); its weakest is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #68 of 69).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#11 / 80↑0.857Source ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#15 / 16↓96Source ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#8 / 16↓57Source ↗official
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct#5 / 10↓65Source ↗official
Anthropic Claude 4 System Cardagentic_coding_safety#2 / 3↑0.88Source ↗official
Anthropic Claude 4 System Cardbenign_request_refusal#1 / 3↓0.0007Source ↗official
Anthropic Claude 4 System Cardharmful_request_safety#3 / 3↑0.9843Source ↗official
Anthropic Claude 4 System Cardstrongreject_jailbreak_success#2 / 3↓0.0714Source ↗official
Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracy#1 / 2↑0.911Source ↗official
Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal#1 / 2↓0.0005Source ↗official
Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety#2 / 2↑0.9727Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#12 / 30↑1195.0Source ↗official
BioTIERpermit_compliance_pct#47 / 52↑88.9Source ↗official
BioTIERrefuse_compliance_pct#6 / 52↑90.7Source ↗official
BullshitBench v2clear_pushback_rate#57 / 122↑0.34Source ↗official
CAIS Risk Indexpolitical_manipulation#40 / 51↓55.3Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#16 / 104↓4.736Source ↗official
Confabulationsconfabulation_rate#1 / 52↓2.723Source ↗official
Emergent Collusionhigh_illegality_game_rate#6 / 13↓0.36Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#29 / 270↑43.41Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#218 / 270↑75Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#21 / 270↑98.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#34 / 268↑99.18Source ↗official
FORTRESSaverage_risk_score#30 / 60↓26.19Source ↗official
FORTRESSover_refusal_score#14 / 59↓2.205Source ↗official
HELM Safetyanthropic_red_team#40 / 80↑0.989Source ↗official
HELM Safetybbq#4 / 80↑0.982Source ↗official
HELM Safetyharmbench#21 / 80↑0.904Source ↗official
HELM Safetysimple_safety_tests#26 / 80↑0.9975Source ↗official
HELM Safetyxstest#24 / 80↑0.97Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#42 / 54↑24.9Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#2 / 69↑94.2Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#24 / 42↓73.5Source ↗official
MASKlying_probability_pct#12 / 53↓15.93Source ↗official
SimpleQA Verifiedf1_score#7 / 13↑28.3Source ↗official
Social Welfare Function Benchmarkfairness#5 / 19↑0.547Source ↗official
SpeciEvalbelief_animal_sentience#97 / 131↑6.58Source ↗official
SpeciEvalland_animal_4ns#53 / 131↓4.42Source ↗official
SpeciEvalsea_animal_4ns#33 / 131↓4.53Source ↗official
SpeciEvalspeciesism#64 / 131↓1.98Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#71 / 94↑88Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#69 / 80↓0.857Source ↗official
Anthropic Claude 4 System Cardbenign_request_refusal#1 / 3↓0.0007Source ↗official
Anthropic Claude 4 System Cardharmful_request_safety#1 / 3↓0.9843Source ↗official
Anthropic Claude 4 System Cardstrongreject_jailbreak_success#2 / 3↑0.0714Source ↗official
Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal#1 / 2↓0.0005Source ↗official
Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety#1 / 2↓0.9727Source ↗official
BioTIERpermit_compliance_pct#47 / 52↑88.9Source ↗official
BioTIERrefuse_compliance_pct#46 / 52↓90.7Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#88 / 104↑4.736Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#53 / 270↓75Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#248 / 270↓98.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#234 / 268↓99.18Source ↗official
FORTRESSaverage_risk_score#31 / 60↑26.19Source ↗official
FORTRESSover_refusal_score#14 / 59↓2.205Source ↗official
HELM Safetyanthropic_red_team#41 / 80↓0.989Source ↗official
HELM Safetyharmbench#60 / 80↓0.904Source ↗official
HELM Safetysimple_safety_tests#54 / 80↓0.9975Source ↗official
HELM Safetyxstest#24 / 80↑0.97Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#68 / 69↓94.2Source ↗official
SpeechMap model completioncomplete_pct#105 / 181↑51.1Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#137 / 156↑0.75Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#137 / 156↑1.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-21.7
Government44.4
Diplomacy67.6
Economy47.2
Society61.2

CAISI CCP narrative alignment

DimensionValueDistribution
CCP narrative alignment3.25