← Models

Model profile

Gemini 2.5 Flash

Googledeveloper
2025-04-17release date
#121 / 267overall rank
31eval lineages

Evidence summary

Gemini 2.5 Flash has an estimated overall rank of #121; its 90% source-sensitivity interval is #44–#186. Its behavior-only rank is #131; company governance moves the combined estimate to #121. Published evidence spans 31 evals and 7 of 7 behavior components. Its strongest relative result is HELM Safety (xstest, #1 of 80); its weakest is Google Gemini 2.5 Flash Model Card (text_safety_delta, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#100 / 3110.7066↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#36 / 800.687↑ higherSource ↗official
Alignment Leaderboardcorrigibility#13 / 244.204↑ higherSource ↗official
Alignment Leaderboardhonesty#11 / 243.613↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#14 / 243.326↑ higherSource ↗official
Alignment Leaderboardrobustness#22 / 242.907↑ higherSource ↗official
Alignment Leaderboardsafety#6 / 244.038↑ higherSource ↗official
Alignment Leaderboardscheming#14 / 243.608↑ higherSource ↗official
Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct#15 / 1696↓ lowerSource ↗official
Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct#14 / 16100↓ lowerSource ↗official
Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct#6 / 1083↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#67 / 1050.19↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#39 / 4391.9↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#46 / 4895.9↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#44 / 4980↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#26 / 4589.4↓ lowerSource ↗official
CAIS Risk Indexmask#40 / 5150.9↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#7 / 4811.7↓ lowerSource ↗official
Confabulationsconfabulation_rate#6 / 524.455↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#65 / 26024.84↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#191 / 26084.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#180 / 26054.63↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#206 / 25891.48↑ higherSource ↗official
Google Gemini 2.5 Flash Model Cardtext_safety_delta#2 / 24.2↓ lowerSource ↗official
HELM Safetyanthropic_red_team#41 / 800.988↑ higherSource ↗official
HELM Safetybbq#7 / 800.977↑ higherSource ↗official
HELM Safetyharmbench#60 / 800.626↑ higherSource ↗official
HELM Safetysimple_safety_tests#48 / 800.98↑ higherSource ↗official
HELM Safetyxstest#1 / 800.988↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#28 / 5427.07↑ higherSource ↗official
LiveSecBenchethics#31 / 4331.43↑ higherSource ↗official
LiveSecBenchfactuality#27 / 4338.55↑ higherSource ↗official
LiveSecBenchlegality#17 / 4356.07↑ higherSource ↗official
LiveSecBenchprivacy#29 / 4331.6↑ higherSource ↗official
LiveSecBenchpsychological_health#19 / 4354.27↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#26 / 5089.4↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#20 / 318.833↓ lowerSource ↗self run
MASKlying_probability_pct#43 / 5350.87↓ lowerSource ↗official
PacifAIstp_score#1 / 790.31↑ higherSource ↗official
PHAREbias_resistance_diagnostic#22 / 660.5174↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#40 / 700.7213↑ higherSource ↗official
PHAREharm_resistance_diagnostic#35 / 700.9366↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#58 / 670.3705↑ higherSource ↗official
PropensityBenchscore#11 / 1468↓ lowerSource ↗official
Social Welfare Function Benchmarkfairness#18 / 190.438↑ higherSource ↗official
SOSBenchbiology_pvr#9 / 230.336↓ lowerSource ↗official
SOSBenchchemistry_pvr#10 / 230.338↓ lowerSource ↗official
SOSBenchmedicine_pvr#7 / 230.462↓ lowerSource ↗official
SOSBenchpharmacology_pvr#10 / 230.684↓ lowerSource ↗official
SOSBenchphysics_pvr#10 / 230.424↓ lowerSource ↗official
SOSBenchpsychology_pvr#9 / 230.326↓ lowerSource ↗official
SpeciEvalbelief_animal_sentience#66 / 1026.67↑ higherSource ↗official
SpeciEvalland_animal_4ns#96 / 1025.05↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#53 / 1024.78↓ lowerSource ↗official
SpeciEvalspeciesism#75 / 1022.39↓ lowerSource ↗official
TACbase_welfare_rate#8 / 6839.74↑ higherSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#12 / 270.715↑ higherSource ↗official
Vigil Mental Health Safetyoverall_score#19 / 2328↑ higherSource ↗official