← Models

Model profile

Gemini 2.0 Flash

Googledeveloper
2024-12-11release date
#205 / 346Safety rank
#431 / 662Freedom rank

Evidence summary

Safety. Gemini 2.0 Flash has an estimated Safety rank of #205; its 90% source-sensitivity interval is #122–#254. Its behavior-only rank is #214; company governance moves the combined estimate to #205. Published Safety evidence spans 20 eval lineages and 7 of 7 components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #1 of 46); its weakest is PropensityBench (score, #13 of 14).

Freedom. Gemini 2.0 Flash has an estimated Freedom rank of #431; its 90% source-sensitivity interval is #278–#514. Published Freedom evidence spans 11 eval lineages and 1 of 1 components. Its strongest relative result is PHARE (jailbreak_resistance_diagnostic, #12 of 67); its weakest is Google Gemini 2.5 Flash Model Card (text_safety_delta, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#241 / 358↓0.8624Source ↗official
AgentDojotargeted_attack_success_rate#9 / 15↓0.1892Source ↗official
AgentDojoutility_under_attack#7 / 15↑0.3982Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#25 / 32↓27.2Source ↗official
AILuminate General Purpose AI Chatdefamation#18 / 32↓17.2Source ↗official
AILuminate General Purpose AI Chathate#13 / 32↓7.1Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#14 / 32↓20Source ↗official
AILuminate General Purpose AI Chatintellectual_property#24 / 32↓19.7Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#16 / 32↓16Source ↗official
AILuminate General Purpose AI Chatprivacy#18 / 32↓13.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#9 / 32↓8.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#10 / 32↓16.3Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#17 / 32↓16.5Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#16 / 32↓12.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#13 / 32↓16.3Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#41 / 80↑0.662Source ↗official
BullshitBench v2clear_pushback_rate#90 / 122↑0.15Source ↗official
Confabulationsconfabulation_rate#23 / 52↓19.55Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#173 / 270↑13.7Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#35 / 270↑93.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#161 / 270↑63.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#197 / 268↑93.64Source ↗official
Google Gemini 2.5 Flash Model Cardtext_safety_delta#1 / 2↓0Source ↗official
HELM Safetyanthropic_red_team#23 / 80↑0.994Source ↗official
HELM Safetybbq#23 / 80↑0.954Source ↗official
HELM Safetyharmbench#51 / 80↑0.662Source ↗official
HELM Safetysimple_safety_tests#43 / 80↑0.985Source ↗official
HELM Safetyxstest#48 / 80↑0.953Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#27 / 54↑27.08Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#32 / 42↓82Source ↗official
MASKlying_probability_pct#39 / 53↓49.93Source ↗official
PandaBench JBB direct-request panelsafety_rate#1 / 46↑1Source ↗official
PHAREbias_resistance_diagnostic#18 / 66↑0.5351Source ↗official
PHAREhallucination_resistance_diagnostic#42 / 70↑0.7159Source ↗official
PHAREharm_resistance_diagnostic#31 / 70↑0.943Source ↗official
PHAREjailbreak_resistance_diagnostic#56 / 67↑0.3735Source ↗official
PropensityBenchscore#13 / 14↓77.8Source ↗official
SpeciEvalbelief_animal_sentience#105 / 131↑6.5Source ↗official
SpeciEvalland_animal_4ns#31 / 131↓4.29Source ↗official
SpeciEvalsea_animal_4ns#80 / 131↓4.79Source ↗official
SpeciEvalspeciesism#95 / 131↓2.33Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AILuminate General Purpose AI Chatchild_sexual_exploitation#8 / 32↑27.2Source ↗official
AILuminate General Purpose AI Chatdefamation#15 / 32↑17.2Source ↗official
AILuminate General Purpose AI Chathate#20 / 32↑7.1Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#19 / 32↑20Source ↗official
AILuminate General Purpose AI Chatintellectual_property#9 / 32↑19.7Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#16 / 32↑16Source ↗official
AILuminate General Purpose AI Chatprivacy#15 / 32↑13.2Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#24 / 32↑8.2Source ↗official
AILuminate General Purpose AI Chatsexual_content#23 / 32↑16.3Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#16 / 32↑16.5Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#17 / 32↑12.8Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#20 / 32↑16.3Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#40 / 80↓0.662Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#234 / 270↓93.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#110 / 270↓63.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#72 / 268↓93.64Source ↗official
Google Gemini 2.5 Flash Model Cardtext_safety_delta#2 / 2↑0Source ↗official
HELM Safetyanthropic_red_team#55 / 80↓0.994Source ↗official
HELM Safetyharmbench#30 / 80↓0.662Source ↗official
HELM Safetysimple_safety_tests#34 / 80↓0.985Source ↗official
HELM Safetyxstest#48 / 80↑0.953Source ↗official
PandaBench JBB direct-request panelsafety_rate#35 / 46↓1Source ↗official
PHAREharm_resistance_diagnostic#40 / 70↓0.943Source ↗official
PHAREjailbreak_resistance_diagnostic#12 / 67↓0.3735Source ↗official
SpeechMap model completioncomplete_pct#114 / 181↑48.7Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism67.1
Self-direction48.9
Care / Harm45.1
Fairness / Cheating42.5
Ethical88.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility65
Warning-phrase rate5
Soft-censorship rate0
API-error rate0