← Models

Model profile

Gemini 1.5 Flash

Googledeveloper
2024-05-14release date
#93 / 267overall rank
15eval lineages

Evidence summary

Gemini 1.5 Flash has an estimated overall rank of #93; its 90% source-sensitivity interval is #36–#167. Its behavior-only rank is #97; company governance moves the combined estimate to #93. Published evidence spans 15 evals and 6 of 7 behavior components. Its strongest relative result is HELM Safety (anthropic_red_team, #2 of 80); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #214 of 258).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Adversarial Robustnessscore#4 / 814↓ lowerSource ↗official
Agent-SafetyBenchcompromise_availability#9 / 1630↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#4 / 1648.4↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#5 / 1639.2↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#6 / 1638.8↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#11 / 1682.4↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#7 / 1641.6↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#4 / 1620.8↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#6 / 1632↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#4 / 150.0787↓ lowerSource ↗official
AgentDojoutility_under_attack#11 / 150.333↑ higherSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#31 / 800.7325↑ higherSource ↗official
AnimalHarmBenchscore#3 / 100.05↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#199 / 26012.14↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#162 / 26086.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#170 / 26055.56↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#214 / 25890.59↑ higherSource ↗official
FORTRESSaverage_risk_score#35 / 4950.61↓ lowerSource ↗official
FORTRESSover_refusal_score#22 / 464.45↓ lowerSource ↗official
HELM Safetyanthropic_red_team#2 / 800.999↑ higherSource ↗official
HELM Safetybbq#32 / 800.947↑ higherSource ↗official
HELM Safetyharmbench#35 / 800.8↑ higherSource ↗official
HELM Safetysimple_safety_tests#58 / 800.97↑ higherSource ↗official
HELM Safetyxstest#66 / 800.921↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#23 / 391.116↓ lowerSource ↗official
OR-Benchover_refusal_rate#17 / 2584.3↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#8 / 251.2↓ lowerSource ↗official
SORRY-Benchavg#5 / 510.08↓ lowerSource ↗official