Evaluation profile
Anthropic Claude Opus 4.1 System Card Addendum
3sub-evals
0.317%total index weight
3components
Within-component eval weight: Fairness 1.12% · Benign helpfulness 1.82% · Misuse resistance 0.226%.
Model score (see each sub-eval direction)Predicted score
About this eval
Developer-reported release-time safety behavior.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| bbq_disambiguated_accuracysystem-cards-2026-08-01/anthropic-claude-opus-4-1.csv:bbq_disambiguated_accuracyAccuracy on disambiguated Bias Benchmark for Question Answering prompts. | fairness_nondiscrimination:1.000anthropic-claude-opus-4-1 | Higher is better | 0.112% | Fairness 1.12% |
| benign_request_refusalsystem-cards-2026-08-01/anthropic-claude-opus-4-1.csv:benign_request_refusalRefusal rate on benign requests. | benign_helpfulness_non_overrefusal:1.000anthropic-claude-opus-4-1 | Lower is better | 0.182% | Benign helpfulness 1.82% |
| harmful_request_safetysystem-cards-2026-08-01/anthropic-claude-opus-4-1.csv:harmful_request_safetyRate of harmless responses to harmful requests. | ordinary_harm_misuse_resistance:1.000anthropic-claude-opus-4-1 | Higher is better | 0.0226% | Misuse resistance 0.226% |
bbq_disambiguated_accuracy
Accuracy on disambiguated Bias Benchmark for Question Answering prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4 | 0.911 | official | |
| 2 | claude-opus-4.1 | 0.907 | official |
benign_request_refusal
Refusal rate on benign requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4 | 0.0005 | official | |
| 2 | claude-opus-4.1 | 0.0008 | official |
harmful_request_safety
Rate of harmless responses to harmful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.1 | 0.9876 | official | |
| 2 | claude-opus-4 | 0.9727 | official |