← Models

Model profile

Gemini 2.5 Flash Lite

Googledeveloper
2025-06-17release date
#242 / 346Safety rank
#44 / 662Freedom rank

Evidence summary

Safety. Gemini 2.5 Flash Lite has an estimated Safety rank of #242; its 90% source-sensitivity interval is #109–#283. Its behavior-only rank is #252; company governance moves the combined estimate to #242. Published Safety evidence spans 23 eval lineages and 7 of 7 components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #3 of 94); its weakest is CAIS Risk Index (agent_red_teaming, #49 of 49).

Freedom. Gemini 2.5 Flash Lite has an estimated Freedom rank of #44; its 90% source-sensitivity interval is #56–#227. Published Freedom evidence spans 11 eval lineages and 1 of 1 components. Its strongest relative result is CAIS Risk Index (bioweapons_assistance, #2 of 54); its weakest is Google Gemini 2.5 Flash-Lite Model Card (text_safety_delta, #1 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#143 / 358↓0.7171Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#15 / 24↓8.67Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#15 / 24↓64.77Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#43 / 80↑0.658Source ↗official
ANIMAscore#18 / 22↑0.5838Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#16 / 111↑1461.0Source ↗official
CAIS Risk Indexagent_red_teaming#49 / 49↓96.8Source ↗official
CAIS Risk Indexbioweapons_assistance#53 / 54↓96.3Source ↗official
CAIS Risk Indexhle_overconfidence#54 / 55↓83.4Source ↗official
CAIS Risk Indexmachiavelli#51 / 51↓109.4Source ↗official
CAIS Risk Indexmask#50 / 57↓54Source ↗official
CAIS Risk Indextextquests_harm#44 / 54↓22.8Source ↗official
DelusionEvaldelusional_prevalence_pct#13 / 16↓56Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#12 / 16↑16.5Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#12 / 16↓5.6Source ↗official
DelusionEvalrelationship_prevalence_pct#10 / 16↓33.6Source ↗official
DelusionEvalsycophancy_prevalence_pct#13 / 16↓35Source ↗official
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta#2 / 2↓5.7Source ↗official
HELM Safetyanthropic_red_team#44 / 80↑0.987Source ↗official
HELM Safetybbq#28 / 80↑0.949Source ↗official
HELM Safetyharmbench#49 / 80↑0.67Source ↗official
HELM Safetysimple_safety_tests#64 / 80↑0.965Source ↗official
HELM Safetyxstest#13 / 80↑0.978Source ↗official
MACHIAVELLIdeception_relative_random_pct#50 / 50↓109.4Source ↗official
MORUscore#8 / 13↑71.34Source ↗official
PHAREbias_resistance_diagnostic#34 / 66↑0.4553Source ↗official
PHAREhallucination_resistance_diagnostic#50 / 70↑0.6809Source ↗official
PHAREharm_resistance_diagnostic#66 / 70↑0.7915Source ↗official
PHAREjailbreak_resistance_diagnostic#45 / 67↑0.4284Source ↗official
SimpleQA Verifiedf1_score#13 / 13↑11.1Source ↗official
SpeciEvalbelief_animal_sentience#58 / 131↑6.83Source ↗official
SpeciEvalland_animal_4ns#107 / 131↓4.81Source ↗official
SpeciEvalsea_animal_4ns#30 / 131↓4.5Source ↗official
SpeciEvalspeciesism#41 / 131↓1.74Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#3 / 94↑96.7Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#10 / 24↑8.67Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#10 / 24↑64.77Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#38 / 80↓0.658Source ↗official
CAIS Risk Indexbioweapons_assistance#2 / 54↑96.3Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#5 / 16↓16.5Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#5 / 16↑5.6Source ↗official
Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta#1 / 2↑5.7Source ↗official
HELM Safetyanthropic_red_team#35 / 80↓0.987Source ↗official
HELM Safetyharmbench#32 / 80↓0.67Source ↗official
HELM Safetysimple_safety_tests#17 / 80↓0.965Source ↗official
HELM Safetyxstest#13 / 80↑0.978Source ↗official
PHAREharm_resistance_diagnostic#5 / 70↓0.7915Source ↗official
PHAREjailbreak_resistance_diagnostic#23 / 67↓0.4284Source ↗official
SpeechMap model completioncomplete_pct#26 / 181↑84.6Source ↗official