← Evals

Evaluation profile

Claude 2 model-card safety and alignment evaluations

4sub-evals
0.138%Safety weight
0.6%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (see each sub-eval direction)Predicted score

About this eval

Human-rated harmlessness and honesty, red-team robustness ordering, and combined HHH response recognition.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
hhhclaude-system-cards/claude-2-scoring.csv:hhhMeasures whether a model recognizes the preferred response across combined HHH comparisons.Safety: truthfulness_honesty:0.500|ordinary_harm_misuse_resistance:0.500anthropic-claude-2-system-cardSafety: higher0.0391%—
human_feedback_harmless_eloclaude-system-cards/claude-2-scoring.csv:human_feedback_harmless_eloMeasures which model responses raters judged less harmful in adversarial conversations.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000anthropic-claude-2-system-cardSafety: higher · Freedom: lower0.021%0.3%
human_feedback_honest_eloclaude-system-cards/claude-2-scoring.csv:human_feedback_honest_eloMeasures which model responses raters judged more accurate and factual.Safety: truthfulness_honesty:1.000anthropic-claude-2-system-cardSafety: higher0.0572%—
red_teaming_rankclaude-system-cards/claude-2-scoring.csv:red_teaming_rankRanks the evaluated Claude systems by performance under Anthropic's red-team evaluation.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000anthropic-claude-2-system-cardSafety: lower · Freedom: higher0.021%0.3%

hhh

Measures whether a model recognizes the preferred response across combined HHH comparisons.

RankModelValueRelative performanceProvenance
1claude-20.86official
2claude-1.30.85official
3claude-instant-1.10.84official

human_feedback_harmless_elo

Measures which model responses raters judged less harmful in adversarial conversations.

RankModelValueRelative performanceProvenance
1claude-1.3282official
2claude-2278official
3claude-instant-1.1243official

human_feedback_honest_elo

Measures which model responses raters judged more accurate and factual.

RankModelValueRelative performanceProvenance
1claude-2137official
2claude-1.3102official
3claude-instant-1.168official

red_teaming_rank

Ranks the evaluated Claude systems by performance under Anthropic's red-team evaluation.

RankModelValueRelative performanceProvenance
1claude-21official
2claude-1.32official
3claude-instant-1.13official