← Models

Model profile

OLMo 2 32B Instruct

Ai2developer
2025-03-13release date
#255 / 346Safety rank
#463 / 662Freedom rank

Evidence summary

Safety. OLMo 2 32B Instruct has an estimated Safety rank of #255; its 90% source-sensitivity interval is #59–#300. Its behavior-only rank is #257; company governance moves the combined estimate to #255. Published Safety evidence spans 8 eval lineages and 6 of 7 components. Its strongest relative result is AILuminate General Purpose AI Chat (child_sexual_exploitation, #6 of 32); its weakest is HELM Safety (bbq, #74 of 80).

Freedom. OLMo 2 32B Instruct has an estimated Freedom rank of #463; its 90% source-sensitivity interval is #290–#546. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is HELM Safety (simple_safety_tests, #31 of 80); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #27 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AgentDrive Safety Compliancescr#29 / 48↑76.25Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#6 / 32↓9.1Source ↗official
AILuminate General Purpose AI Chatdefamation#10 / 32↓10.6Source ↗official
AILuminate General Purpose AI Chathate#8 / 32↓4Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#11 / 32↓18.8Source ↗official
AILuminate General Purpose AI Chatintellectual_property#11 / 32↓8.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#9 / 32↓9Source ↗official
AILuminate General Purpose AI Chatprivacy#10 / 32↓8.1Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#11 / 32↓9.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#15 / 32↓21Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#12 / 32↓13.9Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#10 / 32↓9.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#11 / 32↓12.7Source ↗official
HELM Safetyanthropic_red_team#27 / 80↑0.993Source ↗official
HELM Safetybbq#74 / 80↑0.714Source ↗official
HELM Safetyharmbench#31 / 80↑0.841Source ↗official
HELM Safetysimple_safety_tests#48 / 80↑0.98Source ↗official
HELM Safetyxstest#46 / 80↑0.954Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#21 / 27↑0.615Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#27 / 32↑9.1Source ↗official
AILuminate General Purpose AI Chatdefamation#23 / 32↑10.6Source ↗official
AILuminate General Purpose AI Chathate#25 / 32↑4Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#22 / 32↑18.8Source ↗official
AILuminate General Purpose AI Chatintellectual_property#21 / 32↑8.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#24 / 32↑9Source ↗official
AILuminate General Purpose AI Chatprivacy#23 / 32↑8.1Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#22 / 32↑9.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#18 / 32↑21Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#21 / 32↑13.9Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#22 / 32↑9.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#22 / 32↑12.7Source ↗official
HELM Safetyanthropic_red_team#50 / 80↓0.993Source ↗official
HELM Safetyharmbench#49 / 80↓0.841Source ↗official
HELM Safetysimple_safety_tests#31 / 80↓0.98Source ↗official
HELM Safetyxstest#46 / 80↑0.954Source ↗official