← Models

Model profile

Claude 3 Opus

Anthropicdeveloper
2024-03-04release date
#142 / 346Safety rank
#646 / 662Freedom rank

Evidence summary

Safety. Claude 3 Opus has an estimated Safety rank of #142; its 90% source-sensitivity interval is #82–#228. Its behavior-only rank is #158; company governance moves the combined estimate to #142. Published Safety evidence spans 27 eval lineages and 7 of 7 components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is Claude 3 model-card adversarial human-preference evaluations (multimodal_hallucination_rank, #2 of 2).

Freedom. Claude 3 Opus has an estimated Freedom rank of #646; its 90% source-sensitivity interval is #497–#646. Published Freedom evidence spans 16 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is Claude 3.5 Sonnet model-card safety and alignment evaluations (incorrect_refusals_wildchat, #4 of 4).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#3 / 8↓13Source ↗official
Agent-SafetyBenchcompromise_availability#2 / 16↑43.2Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#3 / 16↑60Source ↗official
Agent-SafetyBenchleak_sensitive_information#1 / 16↑60.4Source ↗official
Agent-SafetyBenchphysical_harm#2 / 16↑61.6Source ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16↑100Source ↗official
Agent-SafetyBenchproperty_loss#1 / 16↑60.4Source ↗official
Agent-SafetyBenchspread_unsafe_information#1 / 16↑35.6Source ↗official
Agent-SafetyBenchviolate_law_ethics#1 / 16↑56.8Source ↗official
AgentDojotargeted_attack_success_rate#7 / 15↓0.1129Source ↗official
AgentDojoutility_under_attack#3 / 15↑0.5246Source ↗official
AgentHarmharm_score#6 / 12↓14.4Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#15 / 80↑0.844Source ↗official
ANIMAscore#19 / 22↑0.573Source ↗official
AnimalHarmBenchscore#4 / 10↑0.043Source ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#5 / 16↓51Source ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#11 / 16↓88Source ↗official
CAIS Risk Indexpolitical_manipulation#50 / 51↓63.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#19 / 104↓5.833Source ↗official
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank#3 / 5↓3Source ↗official
Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rank#3 / 5↓3Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank#3 / 5↓3Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank#1 / 5↓1Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_hallucination_rank#2 / 2↓2Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank#2 / 2↓2Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat#3 / 4↑92Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct#3 / 5↑50Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pct#2 / 5↑50Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat#4 / 4↓11.9Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest#2 / 4↓8.3Source ↗official
COMPL-AI AI-Identity Disclosurescore#1 / 14↑1Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#2 / 14↑0.7557Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#1 / 13↑0.8402Source ↗official
Confabulationsconfabulation_rate#34 / 52↓28.22Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#5 / 270↑74.42Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#12 / 270↑95.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 270↑92.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#7 / 268↑99.91Source ↗official
HELM Safetyanthropic_red_team#7 / 80↑0.998Source ↗official
HELM Safetybbq#38 / 80↑0.94Source ↗official
HELM Safetyharmbench#8 / 80↑0.974Source ↗official
HELM Safetysimple_safety_tests#1 / 80↑1Source ↗official
HELM Safetyxstest#65 / 80↑0.925Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#29 / 39↓1.244Source ↗official
MASKlying_probability_pct#17 / 53↓21Source ↗official
MORUscore#10 / 13↑70.7Source ↗official
OR-Benchover_refusal_rate#20 / 25↓91Source ↗official
OR-Benchtoxic_acceptance_rate#10 / 25↓1.9Source ↗official
SORRY-Benchavg#2 / 51↓0.07Source ↗official
SpeciEvalbelief_animal_sentience#131 / 131↑6.02Source ↗official
SpeciEvalland_animal_4ns#39 / 131↓4.35Source ↗official
SpeciEvalsea_animal_4ns#86 / 131↓4.85Source ↗official
SpeciEvalspeciesism#90 / 131↓2.23Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#6 / 8↑13Source ↗official
Agent-SafetyBenchproduce_unsafe_information#14 / 16↓100Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#66 / 80↓0.844Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#86 / 104↑5.833Source ↗official
Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank#3 / 5↑3Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank#3 / 5↓3Source ↗official
Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank#1 / 5↓1Source ↗official
Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank#1 / 2↑2Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat#2 / 4↓92Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct#2 / 5↓50Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat#4 / 4↓11.9Source ↗official
Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest#2 / 4↓8.3Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#13 / 14↓0.7557Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#258 / 270↓95.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#206 / 270↓92.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#261 / 268↓99.91Source ↗official
HELM Safetyanthropic_red_team#72 / 80↓0.998Source ↗official
HELM Safetyharmbench#72 / 80↓0.974Source ↗official
HELM Safetysimple_safety_tests#58 / 80↓1Source ↗official
HELM Safetyxstest#65 / 80↑0.925Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
OR-Benchover_refusal_rate#20 / 25↓91Source ↗official
OR-Benchtoxic_acceptance_rate#16 / 25↑1.9Source ↗official
SORRY-Benchavg#48 / 51↑0.07Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 156↑1.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#153 / 156↑0Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-13.7
Government47.8
Diplomacy63.1
Economy45.2
Society56.8