Evaluation profile
Claude 3 model-card adversarial human-preference evaluations
Weights below are portfolio-specific global index weights.
About this eval
Refusal calibration, adversarial honesty and harmlessness, multimodal hallucination and harmful-response robustness, and discrimination.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| correct_refusals_wildchat_rankclaude-system-cards/claude-3-scoring.csv:correct_refusals_wildchat_rankRanks evaluated systems by how consistently they refuse harmful real-world prompts. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000anthropic-claude-3-system-card | Safety: lower · Freedom: higher | 0.0136% | 0.0774% |
| discrimination_rankclaude-system-cards/claude-3-scoring.csv:discrimination_rankRanks evaluated systems by lower discriminatory behavior in the card's evaluation. | Safety: fairness_nondiscrimination:1.000anthropic-claude-3-system-card | Safety: lower | 0.0662% | — |
| human_feedback_harmlessness_win_rate_pctclaude-system-cards/claude-3-scoring.csv:human_feedback_harmlessness_win_rate_pctMeasures which responses raters judged less harmful under adversarial prompting. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000anthropic-claude-3-system-card | Safety: higher · Freedom: lower | 0.0121% | 0.0692% |
| human_feedback_honesty_win_rate_pctclaude-system-cards/claude-3-scoring.csv:human_feedback_honesty_win_rate_pctMeasures which responses raters judged more honest under attempts to elicit falsehoods. | Safety: truthfulness_honesty:1.000anthropic-claude-3-system-card | Safety: higher | 0.033% | — |
| incorrect_refusals_wildchat_rankclaude-system-cards/claude-3-scoring.csv:incorrect_refusals_wildchat_rankRanks evaluated systems by how rarely they refuse benign real-world prompts. | Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000anthropic-claude-3-system-card | Safety: lower · Freedom: lower | 0.0941% | 0.0774% |
| incorrect_refusals_xstest_rankclaude-system-cards/claude-3-scoring.csv:incorrect_refusals_xstest_rankCompares exaggerated safety behavior on benign prompts designed to resemble unsafe requests. | Safety: benign_helpfulness_non_overrefusal:1.000 · Freedom: user_direction_following:1.000anthropic-claude-3-system-card | Safety: lower · Freedom: lower | 0.0941% | 0.0774% |
| multimodal_hallucination_rankclaude-system-cards/claude-3-scoring.csv:multimodal_hallucination_rankRanks evaluated multimodal systems by how reliably they avoid hallucinated claims. | Safety: truthfulness_honesty:1.000anthropic-claude-3-system-card | Safety: lower | 0.0234% | — |
| multimodal_harmful_response_rankclaude-system-cards/claude-3-scoring.csv:multimodal_harmful_response_rankRanks evaluated multimodal systems by how consistently they avoid harmful responses. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000anthropic-claude-3-system-card | Safety: lower · Freedom: higher | 0.00857% | 0.049% |
correct_refusals_wildchat_rank
Ranks evaluated systems by how consistently they refuse harmful real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2 | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-2.1 | 4 | official | |
| 5 | claude-3-haiku | 5 | official |
discrimination_rank
Ranks evaluated systems by lower discriminatory behavior in the card's evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-3-haiku | 4 | official | |
| 5 | claude-instant-1.2 | 5 | official |
human_feedback_harmlessness_win_rate_pct
Measures which responses raters judged less harmful under adversarial prompting.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-2.1 | 87 | official | |
| 1 | claude-3-sonnet | 87 | official | |
| 3 | claude-instant-1.2 | 85 | official | |
| 4 | claude-2 | 84 | official |
human_feedback_honesty_win_rate_pct
Measures which responses raters judged more honest under attempts to elicit falsehoods.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 69 | official | |
| 2 | claude-2.1 | 68 | official | |
| 3 | claude-2 | 66 | official | |
| 4 | claude-instant-1.2 | 61 | official |
incorrect_refusals_wildchat_rank
Ranks evaluated systems by how rarely they refuse benign real-world prompts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-haiku | 1 | official | |
| 2 | claude-3-sonnet | 2 | official | |
| 3 | claude-3-opus | 3 | official | |
| 4 | claude-2 | 4 | official | |
| 5 | claude-2.1 | 5 | official |
incorrect_refusals_xstest_rank
Compares exaggerated safety behavior on benign prompts designed to resemble unsafe requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 1 | official | |
| 2 | claude-2 | 2 | official | |
| 3 | claude-3-haiku | 3 | official | |
| 4 | claude-2.1 | 4 | official | |
| 4 | claude-3-sonnet | 4 | official |
multimodal_hallucination_rank
Ranks evaluated multimodal systems by how reliably they avoid hallucinated claims.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 1 | official | |
| 2 | claude-3-opus | 2 | official |
multimodal_harmful_response_rank
Ranks evaluated multimodal systems by how consistently they avoid harmful responses.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-sonnet | 1 | official | |
| 2 | claude-3-opus | 2 | official |