← Models

Model profile

Claude 3.5 Haiku

Anthropicdeveloper
2024-10-22release date
#77 / 346Safety rank
#639 / 662Freedom rank

Evidence summary

Safety. Claude 3.5 Haiku has an estimated Safety rank of #77; its 90% source-sensitivity interval is #16–#205. Its behavior-only rank is #88; company governance moves the combined estimate to #77. Published Safety evidence spans 15 eval lineages and 7 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #7 of 270); its weakest is Confabulations (confabulation_rate, #51 of 52).

Freedom. Claude 3.5 Haiku has an estimated Freedom rank of #639; its 90% source-sensitivity interval is #467–#646. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is Anthropic Claude Haiku 4.5 System Card (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#63 / 358↓0.4115Source ↗official
Agent-SafetyBenchcompromise_availability#11 / 16↑26.4Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#2 / 16↑60.8Source ↗official
Agent-SafetyBenchleak_sensitive_information#3 / 16↑47.2Source ↗official
Agent-SafetyBenchphysical_harm#4 / 16↑45.6Source ↗official
Agent-SafetyBenchproduce_unsafe_information#1 / 16↑100Source ↗official
Agent-SafetyBenchproperty_loss#5 / 16↑46Source ↗official
Agent-SafetyBenchspread_unsafe_information#2 / 16↑33.6Source ↗official
Agent-SafetyBenchviolate_law_ethics#3 / 16↑41.2Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#2 / 32↓1.8Source ↗official
AILuminate General Purpose AI Chatdefamation#3 / 32↓2.5Source ↗official
AILuminate General Purpose AI Chathate#3 / 32↓1Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#3 / 32↓3.3Source ↗official
AILuminate General Purpose AI Chatintellectual_property#3 / 32↓2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#2 / 32↓3Source ↗official
AILuminate General Purpose AI Chatprivacy#3 / 32↓2.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#2 / 32↓3Source ↗official
AILuminate General Purpose AI Chatsexual_content#3 / 32↓2.503Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#3 / 32↓3.715Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#3 / 32↓2.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#2 / 32↓2.9Source ↗official
AnimalHarmBenchscore#6 / 10↑0.02Source ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#3 / 16↓10Source ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#4 / 16↓19Source ↗official
Anthropic Claude Haiku 4.5 System Cardagentic_coding_safety#1 / 3↑1Source ↗official
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal#1 / 3↑0.7Source ↗official
Anthropic Claude Haiku 4.5 System Cardharmful_request_safety#1 / 2↑0.9972Source ↗official
BullshitBench v2clear_pushback_rate#36 / 122↑0.5Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#36 / 104↓20.74Source ↗official
Confabulationsconfabulation_rate#51 / 52↓65.84Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#14 / 270↑56.59Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#7 / 270↑97Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#50 / 270↑94.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#28 / 268↑99.36Source ↗official
FORTRESSaverage_risk_score#37 / 60↓30.41Source ↗official
FORTRESSover_refusal_score#54 / 59↓13.16Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#24 / 39↓1.118Source ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 46↑0.99Source ↗official
PHAREbias_resistance_diagnostic#49 / 66↑0.3808Source ↗official
PHAREhallucination_resistance_diagnostic#27 / 70↑0.7804Source ↗official
PHAREharm_resistance_diagnostic#25 / 70↑0.9536Source ↗official
PHAREjailbreak_resistance_diagnostic#21 / 67↑0.6482Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Agent-SafetyBenchproduce_unsafe_information#14 / 16↓100Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#31 / 32↑1.8Source ↗official
AILuminate General Purpose AI Chatdefamation#30 / 32↑2.5Source ↗official
AILuminate General Purpose AI Chathate#30 / 32↑1Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#30 / 32↑3.3Source ↗official
AILuminate General Purpose AI Chatintellectual_property#30 / 32↑2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#31 / 32↑3Source ↗official
AILuminate General Purpose AI Chatprivacy#30 / 32↑2.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#31 / 32↑3Source ↗official
AILuminate General Purpose AI Chatsexual_content#30 / 32↑2.503Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#30 / 32↑3.715Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#30 / 32↑2.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#31 / 32↑2.9Source ↗official
Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal#3 / 3↓0.7Source ↗official
Anthropic Claude Haiku 4.5 System Cardharmful_request_safety#2 / 2↓0.9972Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#69 / 104↑20.74Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#264 / 270↓97Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#217 / 270↓94.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#239 / 268↓99.36Source ↗official
FORTRESSaverage_risk_score#24 / 60↑30.41Source ↗official
FORTRESSover_refusal_score#54 / 59↓13.16Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
PandaBench JBB direct-request panelsafety_rate#24 / 46↓0.99Source ↗official
PHAREharm_resistance_diagnostic#46 / 70↓0.9536Source ↗official
PHAREjailbreak_resistance_diagnostic#47 / 67↓0.6482Source ↗official
SpeechMap model completioncomplete_pct#181 / 181↑3.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism70.1
Self-direction50.6
Care / Harm55.9
Fairness / Cheating54.7
Ethical90.4