Model profile
Evidence summary
Safety. Claude Sonnet 5.5 has an estimated Safety rank of #15; its 90% source-sensitivity interval is #3–#81. Its behavior-only rank is #23; company governance moves the combined estimate to #15. Published Safety evidence spans 5 eval lineages and 4 of 7 components. Its strongest relative result is SM-Bench (ambiguous_interpretation, #2 of 92); its weakest is Olam Diplomacy — Broken Promise Rate (broken_promise_rate_pct, #15 of 19).
Freedom. Claude Sonnet 5.5 has an estimated Freedom rank of #442; its 90% source-sensitivity interval is #255–#556. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (overfit, #20 of 92); its weakest is SM-Bench (adversarial, #67 of 92).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗claude-sonnet-5-5
- OpenRouter ↗anthropic/claude-sonnet-5.5
- Official launch announcement ↗Exact model document · anthropic · first party
Safety evals
Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Source |
|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | #76 / 358 | ↓0.4701 | Source ↗official | |
| BullshitBench v2clear_pushback_rate | #25 / 122 | ↑0.59 | Source ↗official | |
| Olam Diplomacy — Broken Promise Ratebroken_promise_rate_pct | #15 / 19 | ↓17.7 | Source ↗official | |
| SM-Benchadversarial | #24 / 92 | ↑85.37 | Source ↗official | |
| SM-Benchambiguous_interpretation | #2 / 92 | ↑97.02 | Source ↗official | |
| SM-Benchanti_hallucination | #10 / 92 | ↑98.95 | Source ↗official | |
| SM-Bencheq_boundaries | #60 / 92 | ↑61.52 | Source ↗official | |
| SM-Benchoverfit | #20 / 92 | ↑89.62 | Source ↗official | |
| SpeciEvalbelief_animal_sentience | #78 / 131 | ↑6.75 | Source ↗official | |
| SpeciEvalland_animal_4ns | #36 / 131 | ↓4.33 | Source ↗official | |
| SpeciEvalsea_animal_4ns | #93 / 131 | ↓4.9 | Source ↗official | |
| SpeciEvalspeciesism | #48 / 131 | ↓1.82 | Source ↗official |
Freedom evals
Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.