← Models

Model profile

Gemini 3.7 Flash

Googledeveloper
2026-08-13release date
#24 / 346Safety rank
#396 / 662Freedom rank

Evidence summary

Safety. Gemini 3.7 Flash has an estimated Safety rank of #24; its 90% source-sensitivity interval is #12–#102. Its behavior-only rank is #29; company governance moves the combined estimate to #24. Published Safety evidence spans 19 eval lineages and 6 of 7 components. Its strongest relative result is Manager Coercion Bench (fabrication_rate, #1 of 15); its weakest is SpeciEval (speciesism, #120 of 131).

Freedom. Gemini 3.7 Flash has an estimated Freedom rank of #396; its 90% source-sensitivity interval is #155–#554. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (overfit, #8 of 92); its weakest is SM-Bench (adversarial, #77 of 92).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#120 / 358↓0.6453Source ↗official
BioSecBench-Refusal V2redteam_refusal_pct#5 / 6↑63.64Source ↗official
BioSecBench-Refusal V2routine_compliance_pct#3 / 6↑49.72Source ↗official
BullshitBench v2clear_pushback_rate#62 / 122↑0.33Source ↗official
CAIS Risk Indexagent_red_teaming#4 / 49↓38.5Source ↗official
CAIS Risk Indexbioweapons_assistance#27 / 54↓65.9Source ↗official
CAIS Risk Indexhle_overconfidence#26 / 55↓51.9Source ↗official
CAIS Risk Indexmachiavelli#21 / 51↓85.5Source ↗official
CAIS Risk Indexmask#43 / 57↓46.3Source ↗official
CAIS Risk Indexpolitical_manipulation#22 / 51↓45.1Source ↗official
CAIS Risk Indextextquests_harm#49 / 54↓24.4Source ↗official
Claude system cards — Gray Swan Q1+Q2 indirect prompt injection k=15attack_success_probability_k15_pct#6 / 12↓9.2Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#7 / 15↓9.2Source ↗official
Manager Coercion Benchcoercion_ladder_depth#30 / 45↓8.7Source ↗official
Manager Coercion Benchfabrication_rate#1 / 15↓0Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#10 / 24↓7Source ↗official
Pander Scoreconversational_absolute_pander_score#16 / 26↓15.68Source ↗official
Pander Scoreinstructional_absolute_pander_score#18 / 26↓65.05Source ↗official
SM-Benchadversarial#13 / 92↑86.83Source ↗official
SM-Benchambiguous_interpretation#17 / 92↑90.48Source ↗official
SM-Benchanti_hallucination#27 / 92↑97.38Source ↗official
SM-Bencheq_boundaries#42 / 92↑65.73Source ↗official
SM-Benchoverfit#8 / 92↑93.99Source ↗official
SpeciEvalbelief_animal_sentience#49 / 131↑6.87Source ↗official
SpeciEvalland_animal_4ns#14 / 131↓4.17Source ↗official
SpeciEvalsea_animal_4ns#54 / 131↓4.67Source ↗official
SpeciEvalspeciesism#120 / 131↓2.77Source ↗official
TACbase_welfare_rate#52 / 92↑25.64Source ↗official
Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct#1 / 14↓0Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
BioSecBench-Refusal V2redteam_refusal_pct#2 / 6↓63.64Source ↗official
BioSecBench-Refusal V2routine_compliance_pct#3 / 6↑49.72Source ↗official
CAIS Risk Indexbioweapons_assistance#28 / 54↑65.9Source ↗official
SM-Benchadversarial#77 / 92↓86.83Source ↗official
SM-Bencheq_boundaries#42 / 92↑65.73Source ↗official
SM-Benchoverfit#8 / 92↑93.99Source ↗official