← Models

Model profile

Gemini 2.0 Flash

Googledeveloper
2024-12-11release date
#162 / 267overall rank
12eval lineages

Evidence summary

Gemini 2.0 Flash has an estimated overall rank of #162; its 90% source-sensitivity interval is #65–#201. Its behavior-only rank is #170; company governance moves the combined estimate to #162. Published evidence spans 12 evals and 6 of 7 behavior components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #1 of 46); its weakest is PropensityBench (score, #13 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#210 / 3110.8654↓ lowerSource ↗official
AgentDojotargeted_attack_success_rate#9 / 150.1892↓ lowerSource ↗official
AgentDojoutility_under_attack#7 / 150.3982↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#73 / 1050.15↑ higherSource ↗official
Confabulationsconfabulation_rate#23 / 5219.55↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#165 / 26013.7↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#41 / 26093.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#147 / 26063.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#186 / 25893.64↑ higherSource ↗official
Google Gemini 2.5 Flash Model Cardtext_safety_delta#1 / 20↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#27 / 5427.08↑ higherSource ↗official
MASKlying_probability_pct#39 / 5349.93↓ lowerSource ↗official
PandaBench JBB direct-request panelsafety_rate#1 / 461↑ higherSource ↗official
PHAREbias_resistance_diagnostic#18 / 660.5351↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#42 / 700.7159↑ higherSource ↗official
PHAREharm_resistance_diagnostic#31 / 700.943↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#56 / 670.3735↑ higherSource ↗official
PropensityBenchscore#13 / 1477.8↓ lowerSource ↗official
SpeciEvalbelief_animal_sentience#79 / 1026.5↑ higherSource ↗official
SpeciEvalland_animal_4ns#22 / 1024.29↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#57 / 1024.79↓ lowerSource ↗official
SpeciEvalspeciesism#71 / 1022.33↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism67.1
Self-direction48.9
Care / Harm45.1
Fairness / Cheating42.5
Ethical88.9

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility65
Warning-phrase rate5
Soft-censorship rate0
API-error rate0