← Models

Model profile

Claude Opus 4.8

Anthropicdeveloper
2026-05-28release date
#4 / 267overall rank
20eval lineages
3discovery sources

Evidence summary

Claude Opus 4.8 has an estimated overall rank of #4; its 90% source-sensitivity interval is #3–#49. Its behavior-only rank is #5; company governance moves the combined estimate to #4. Published evidence spans 20 evals and 7 of 7 behavior components. Its strongest relative result is BullshitBench v2 (clear_pushback_rate, #1 of 105); its weakest is SpeciEval (belief_animal_sentience, #89 of 102).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#25 / 3110.3589↓ lowerSource ↗official
BioSecBench-Refusalbalanced_refusal_score#7 / 100.3633↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#1 / 1050.945↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#8 / 4347↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#5 / 4826.4↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#3 / 4926.4↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#10 / 4583.7↓ lowerSource ↗official
CAIS Risk Indexmask#9 / 518.1↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#26 / 3253.8↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#35 / 4821↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#4 / 1052.293↓ lowerSource ↗official
DystopiaBenchbasaglia_score#1 / 5018.7↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#1 / 5021.27↓ lowerSource ↗official
DystopiaBenchhuxley_score#1 / 5010.63↓ lowerSource ↗official
DystopiaBenchlaguardia_score#1 / 5022↓ lowerSource ↗official
DystopiaBenchorwell_score#1 / 5017.6↓ lowerSource ↗official
DystopiaBenchpetrov_score#2 / 5024.2↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#16 / 26051.94↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#54 / 26092↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#7 / 26099.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#39 / 25899.27↑ higherSource ↗official
FORTRESSaverage_risk_score#16 / 4918.19↓ lowerSource ↗official
FORTRESSover_refusal_score#34 / 467.42↓ lowerSource ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#4 / 135.5↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#30 / 5426.79↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#10 / 5083.7↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#4 / 314.5↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
SM-Benchadversarial#42 / 7380.49↑ higherSource ↗official
SM-Benchambiguous_interpretation#11 / 7390.48↑ higherSource ↗official
SM-Benchanti_hallucination#1 / 73100↑ higherSource ↗official
SM-Bencheq_boundaries#61 / 7353.65↑ higherSource ↗official
SM-Benchoverfit#14 / 7387.43↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#89 / 1026.38↑ higherSource ↗official
SpeciEvalland_animal_4ns#35 / 1024.4↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#41 / 1024.7↓ lowerSource ↗official
SpeciEvalspeciesism#76 / 1022.4↓ lowerSource ↗official
TACbase_welfare_rate#1 / 6864.74↑ higherSource ↗official
UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pct#2 / 50.8↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-15.2
Government45.2
Diplomacy61.8
Economy46.3
Society56.6