← Models

Model profile

Claude Opus 5

Anthropicdeveloper
2026-07-24release date
#1 / 267overall rank
7eval lineages
1discovery sources

Evidence summary

Claude Opus 5 has an estimated overall rank of #1; its 90% source-sensitivity interval is #1–#6. Published evidence spans 7 evals and 5 of 7 behavior components. Its strongest relative result is SM-Bench (ambiguous_interpretation, #1 of 73); its weakest is SpeciEval (belief_animal_sentience, #79 of 102).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#52 / 3110.5007↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#11 / 1050.715↑ higherSource ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#1 / 132↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#2 / 313.5↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
SM-Benchadversarial#16 / 7386.34↑ higherSource ↗official
SM-Benchambiguous_interpretation#1 / 7397.62↑ higherSource ↗official
SM-Benchanti_hallucination#7 / 7398.95↑ higherSource ↗official
SM-Bencheq_boundaries#21 / 7368.82↑ higherSource ↗official
SM-Benchoverfit#11 / 7391.26↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#79 / 1026.5↑ higherSource ↗official
SpeciEvalland_animal_4ns#27 / 1024.35↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#13 / 1024.38↓ lowerSource ↗official
SpeciEvalspeciesism#43 / 1021.92↓ lowerSource ↗official
TACbase_welfare_rate#2 / 6859.62↑ higherSource ↗official
UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pct#1 / 50.1↓ lowerSource ↗official