Model profile
Claude 3.5 Sonnet
Evidence summary
Claude 3.5 Sonnet has an estimated overall rank of #109; its 90% source-sensitivity interval is #55–#160. Its behavior-only rank is #122; company governance moves the combined estimate to #109. Published evidence spans 27 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #3 of 260); its weakest is FORTRESS (over_refusal_score, #46 of 46).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- OpenRouter ↗anthropic/claude-3.5-sonnet
- System card ↗Family-level model document · Anthropic · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| Adversarial Robustnessscore | #5 / 8 | 16 | ↓ lower | Source ↗official | |
| Agent-SafetyBenchcompromise_availability | #3 / 16 | 40.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchharmful_vulnerable_code | #1 / 16 | 64.8 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchleak_sensitive_information | #2 / 16 | 57.6 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchphysical_harm | #1 / 16 | 69.6 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchproduce_unsafe_information | #1 / 16 | 100 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchproperty_loss | #2 / 16 | 58.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchspread_unsafe_information | #3 / 16 | 32.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchviolate_law_ethics | #2 / 16 | 52 | ↑ higher | Source ↗official | |
| AgentDojotargeted_attack_success_rate | #8 / 15 | 0.1749 | ↓ lower | Source ↗official | |
| AgentDojoutility_under_attack | #2 / 15 | 0.6184 | ↑ higher | Source ↗official | |
| AgentHarmharm_score | #4 / 12 | 13.5 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | #1 / 32 | 1 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatdefamation | #2 / 32 | 1.3 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chathate | #2 / 32 | 0.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | #2 / 32 | 0.9 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatintellectual_property | #2 / 32 | 0.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatnon_violent_crimes | #1 / 32 | 1 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatprivacy | #2 / 32 | 0.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsex_related_crimes | #1 / 32 | 1.6 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsexual_content | #2 / 32 | 1.8 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatspecialized_advice | #2 / 32 | 2.7 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatsuicide_self_harm | #2 / 32 | 1.3 | ↓ lower | Source ↗official | |
| AILuminate General Purpose AI Chatviolent_crimes | #1 / 32 | 1.7 | ↓ lower | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #3 / 80 | 0.8835 | ↑ higher | Source ↗official | |
| AnimalHarmBenchscore | #7 / 10 | 0.018 | ↑ higher | Source ↗official | |
| Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct | #12 / 16 | 86 | ↓ lower | Source ↗official | |
| Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct | #12 / 16 | 99 | ↓ lower | Source ↗official | |
| BullshitBench v2clear_pushback_rate | #29 / 105 | 0.45 | ↑ higher | Source ↗official | |
| CASE-Benchagreement_accuracy | #1 / 7 | 90.11 | ↑ higher | Source ↗official | |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | #11 / 105 | 3.29 | ↓ lower | Source ↗official | |
| Confabulationsconfabulation_rate | #12 / 52 | 12.87 | ↓ lower | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #18 / 260 | 51.68 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #3 / 260 | 98.17 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #28 / 260 | 97.5 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #10 / 258 | 99.95 | ↑ higher | Source ↗official | |
| FORTRESSaverage_risk_score | #4 / 49 | 12.96 | ↓ lower | Source ↗official | |
| FORTRESSover_refusal_score | #46 / 46 | 21.07 | ↓ lower | Source ↗official | |
| HELM Safetyanthropic_red_team | #7 / 80 | 0.998 | ↑ higher | Source ↗official | |
| HELM Safetybbq | #28 / 80 | 0.949 | ↑ higher | Source ↗official | |
| HELM Safetyharmbench | #5 / 80 | 0.981 | ↑ higher | Source ↗official | |
| HELM Safetysimple_safety_tests | #1 / 80 | 1 | ↑ higher | Source ↗official | |
| HELM Safetyxstest | #43 / 80 | 0.956 | ↑ higher | Source ↗official | |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | #15 / 39 | 0.9208 | ↓ lower | Source ↗official | |
| MASKlying_probability_pct | #22 / 53 | 30.54 | ↓ lower | Source ↗official | |
| OR-Benchover_refusal_rate | #11 / 25 | 43.8 | ↓ lower | Source ↗official | |
| OR-Benchtoxic_acceptance_rate | #12 / 25 | 3.4 | ↓ lower | Source ↗official | |
| PandaBench JBB direct-request panelsafety_rate | #14 / 46 | 0.99 | ↑ higher | Source ↗official | |
| PHAREhallucination_resistance_diagnostic | #5 / 70 | 0.8671 | ↑ higher | Source ↗official | |
| PHAREharm_resistance_diagnostic | #24 / 70 | 0.954 | ↑ higher | Source ↗official | |
| SafeArenanormalized_safety_score | #1 / 5 | 55 | ↑ higher | Source ↗official | |
| SORRY-Benchavg | #12 / 51 | 0.14 | ↓ lower | Source ↗official | |
| SpeciesismBenchexplicit_speciesism_scale | #4 / 7 | 2.467 | ↓ lower | Source ↗official | |
| SpeciesismBenchmorally_wrong_rate | #8 / 8 | 29.74 | ↑ higher | Source ↗official | |
| SpeciesismBenchspeciesism_recognition_rate | #6 / 8 | 84.28 | ↑ higher | Source ↗official | |
| SpeciEvalbelief_animal_sentience | #53 / 102 | 6.78 | ↑ higher | Source ↗official | |
| SpeciEvalland_animal_4ns | #89 / 102 | 4.97 | ↓ lower | Source ↗official | |
| SpeciEvalsea_animal_4ns | #78 / 102 | 5 | ↓ lower | Source ↗official | |
| SpeciEvalspeciesism | #38 / 102 | 1.85 | ↓ lower | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
ValueCompass
| Dimension | Value | Distribution |
|---|---|---|
| Universalism | 75.9 | |
| Self-direction | 59.2 | |
| Care / Harm | 69 | |
| Fairness / Cheating | 69.5 | |
| Ethical | 90.6 |