← Models

Model profile

OLMo 7B Instruct

Ai2developer
2024-07-09release date
Not rankedSafety rank
#115 / 662Freedom rank

Evidence summary

Safety. OLMo 7B Instruct does not meet the evidence gate for a Safety rank. Published Safety evidence spans 2 eval lineages and 5 of 7 components. Its strongest relative result is AbstentionBench (stale_f1, #10 of 20); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #32 of 32).

Freedom. OLMo 7B Instruct has an estimated Freedom rank of #115; its 90% source-sensitivity interval is #1–#555. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is AILuminate General Purpose AI Chat (child_sexual_exploitation, #1 of 32); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #1 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AbstentionBenchanswer_unknown_f1#16 / 20↑0.8086Source ↗official
AbstentionBenchfalse_premise_f1#13 / 20↑0.6608Source ↗official
AbstentionBenchstale_f1#10 / 20↑0.6392Source ↗official
AbstentionBenchsubjective_f1#16 / 20↑0.6598Source ↗official
AbstentionBenchunderspecified_context_f1#15 / 20↑0.5001Source ↗official
AbstentionBenchunderspecified_intent_f1#16 / 20↑0.6654Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#32 / 32↓81.2Source ↗official
AILuminate General Purpose AI Chatdefamation#32 / 32↓74Source ↗official
AILuminate General Purpose AI Chathate#32 / 32↓65.6Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#32 / 32↓79.2Source ↗official
AILuminate General Purpose AI Chatintellectual_property#32 / 32↓53.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#32 / 32↓75Source ↗official
AILuminate General Purpose AI Chatprivacy#32 / 32↓62.8Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#32 / 32↓75.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#32 / 32↓70.5Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#32 / 32↓63.4Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#32 / 32↓70.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#32 / 32↓77.6Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#1 / 32↑81.2Source ↗official
AILuminate General Purpose AI Chatdefamation#1 / 32↑74Source ↗official
AILuminate General Purpose AI Chathate#1 / 32↑65.6Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#1 / 32↑79.2Source ↗official
AILuminate General Purpose AI Chatintellectual_property#1 / 32↑53.2Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#1 / 32↑75Source ↗official
AILuminate General Purpose AI Chatprivacy#1 / 32↑62.8Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#1 / 32↑75.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#1 / 32↑70.5Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#1 / 32↑63.4Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#1 / 32↑70.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#1 / 32↑77.6Source ↗official