← Models

Model profile

Gemini 2 Flash

Googledeveloper
2024-12-11release date
#130 / 267overall rank
7eval lineages

Evidence summary

Gemini 2 Flash has an estimated overall rank of #130; its 90% source-sensitivity interval is #78–#194. Its behavior-only rank is #141; company governance moves the combined estimate to #130. Published evidence spans 7 evals and 5 of 7 behavior components. Its strongest relative result is AILuminate General Purpose AI Chat (sex_related_crimes, #9 of 32); its weakest is AILuminate General Purpose AI Chat (child_sexual_exploitation, #25 of 32).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#25 / 3227.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#18 / 3217.2↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#13 / 327.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#14 / 3220↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#24 / 3219.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#16 / 3216↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#18 / 3213.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#9 / 328.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#10 / 3216.3↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#17 / 3216.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#16 / 3212.8↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#13 / 3216.3↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#41 / 800.662↑ higherSource ↗official
HELM Safetyanthropic_red_team#23 / 800.994↑ higherSource ↗official
HELM Safetybbq#23 / 800.954↑ higherSource ↗official
HELM Safetyharmbench#51 / 800.662↑ higherSource ↗official
HELM Safetysimple_safety_tests#43 / 800.985↑ higherSource ↗official
HELM Safetyxstest#48 / 800.953↑ higherSource ↗official