← Models

Model profile

Claude 2

Anthropicdeveloper
2023-07-11release date
#3 / 267overall rank
6eval lineages

Evidence summary

Claude 2 has an estimated overall rank of #3; its 90% source-sensitivity interval is #1–#36. Its behavior-only rank is #4; company governance moves the combined estimate to #3. Published evidence spans 6 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #1 of 33); its weakest is SALAD-Bench (mcq_misinformation_harms, #23 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
DecodingTrustmachine_ethics#3 / 885.17↑ higherSource ↗official
DecodingTruststereotype_bias#1 / 8100↑ higherSource ↗official
DecodingTrusttoxicity#1 / 892.11↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#2 / 1485.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#3 / 1498.67↑ higherSource ↗official
HarmBenchdr#2 / 282↓ lowerSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#1 / 3386.64↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#1 / 3393.49↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#1 / 3387.03↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#1 / 3391.61↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#1 / 3387.15↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#1 / 3388.31↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#1 / 3399.88↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#2 / 3399.66↑ higherSource ↗official
SALAD-Benchbase_malicious_use#1 / 3399.97↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#1 / 3399.7↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#1 / 3399.58↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#1 / 3399.41↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#22 / 3321.67↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#20 / 3328.61↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#23 / 3319.17↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#23 / 3320.71↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#21 / 3324.38↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#20 / 3329.44↑ higherSource ↗official
SORRY-Benchavg#2 / 510.07↓ lowerSource ↗official
SuperCLUE Safetyinstruction_attack#8 / 3170.69↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#1 / 3178.18↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#14 / 3177.66↑ higherSource ↗official