← Models

Model profile

Gemini 3 Flash Preview

Googledeveloper
2025-12-17release date
#219 / 346Safety rank
#346 / 662Freedom rank

Evidence summary

Safety. Gemini 3 Flash Preview has an estimated Safety rank of #219; its 90% source-sensitivity interval is #125–#268. Its behavior-only rank is #234; company governance moves the combined estimate to #219. Published Safety evidence spans 29 eval lineages and 7 of 7 components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 131); its weakest is Adversarial Humanities Benchmark (AHB) — Table 5 (child_sexual_exploitation, #31 of 31).

Freedom. Gemini 3 Flash Preview has an estimated Freedom rank of #346; its 90% source-sensitivity interval is #159–#508. Published Freedom evidence spans 8 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is BioTIER (permit_compliance_pct, #38 of 52).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#316 / 358↓0.9241Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#31 / 31↓83.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#28 / 31↓84.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#24 / 31↓66.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#29 / 31↓88.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#31 / 31↓81.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#29 / 31↓93.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#30 / 31↓85.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#31 / 31↓93.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#28 / 31↓66.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#30 / 31↓76.53Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#28 / 31↓77.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#27 / 31↓88.9Source ↗official
AgentAbstainabstain#16 / 17↑43.6Source ↗official
AgentAbstaincar#15 / 17↑43.4Source ↗official
AgentAbstainpaired#14 / 17↑39.7Source ↗official
Alignment Leaderboardcorrigibility#10 / 24↑4.32Source ↗official
Alignment Leaderboardhonesty#9 / 24↑3.682Source ↗official
Alignment Leaderboardnon_manipulation#12 / 24↑3.51Source ↗official
Alignment Leaderboardrobustness#14 / 24↑3.373Source ↗official
Alignment Leaderboardsafety#11 / 24↑3.875Source ↗official
Alignment Leaderboardscheming#12 / 24↑3.645Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#25 / 30↑1153.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#32 / 111↑1449.0Source ↗official
AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent#3 / 17↓21.8Source ↗official
AuAu Authoritarian Response Auditrealistic_prompt_arr_percent#10 / 17↓2.2Source ↗official
BioTIERpermit_compliance_pct#38 / 52↑98.3Source ↗official
BioTIERrefuse_compliance_pct#22 / 52↑54.3Source ↗official
BullshitBench v2clear_pushback_rate#102 / 122↑0.1Source ↗official
CAIS Risk Indexagent_red_teaming#47 / 49↓92.9Source ↗official
CAIS Risk Indexbioweapons_assistance#35 / 54↓75.3Source ↗official
CAIS Risk Indexhle_overconfidence#28 / 55↓55.1Source ↗official
CAIS Risk Indexmachiavelli#42 / 51↓95.3Source ↗official
CAIS Risk Indexmask#54 / 57↓61.7Source ↗official
CAIS Risk Indextextquests_harm#25 / 54↓17.8Source ↗official
DystopiaBenchbasaglia_score#47 / 50↓73.2Source ↗official
DystopiaBenchbaudrillard_score#50 / 50↓81.13Source ↗official
DystopiaBenchhuxley_score#47 / 50↓83Source ↗official
DystopiaBenchlaguardia_score#46 / 50↓71.63Source ↗official
DystopiaBenchorwell_score#31 / 50↓72.43Source ↗official
DystopiaBenchpetrov_score#43 / 50↓80.33Source ↗official
Every Model Cheats — Cybench Cheat Propensitybaseline_cheat_propensity_pct#10 / 22↓30.4Source ↗official
Every Model Cheats — Cybench Cheat Propensitysevere_anti_cheat_cheat_propensity_pct#1 / 22↓0Source ↗official
Every Model Cheats — Cybench Cheat Propensitystandard_anti_cheat_cheat_propensity_pct#22 / 22↓69.6Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
JuICE Cultural-Error Span Detectionf1#3 / 10↑0.4839Source ↗official
MACHIAVELLIdeception_relative_random_pct#39 / 50↓95.3Source ↗official
MT-JailBench CrescendoXsafety_score#16 / 21↑13.84Source ↗official
NESSiE Necessary Safety Benchmarksh_score#11 / 16↑79.49Source ↗official
Pander Scoreconversational_absolute_pander_score#24 / 26↓25.97Source ↗official
Pander Scoreinstructional_absolute_pander_score#24 / 26↓70.12Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#12 / 17↑0.227Source ↗official
SM-Benchadversarial#28 / 92↑84.39Source ↗official
SM-Benchambiguous_interpretation#69 / 92↑80.65Source ↗official
SM-Benchanti_hallucination#78 / 92↑82.72Source ↗official
SM-Bencheq_boundaries#43 / 92↑65.45Source ↗official
SM-Benchoverfit#8 / 92↑93.99Source ↗official
SpeciEvalbelief_animal_sentience#1 / 131↑7Source ↗official
SpeciEvalland_animal_4ns#92 / 131↓4.72Source ↗official
SpeciEvalsea_animal_4ns#105 / 131↓5.03Source ↗official
SpeciEvalspeciesism#131 / 131↓3.9Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#23 / 23↑82.56Source ↗official
TACbase_welfare_rate#19 / 92↑35.9Source ↗self run
Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct#2 / 14↓0.3745Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#79 / 94↑86.5Source ↗official
Vigil Mental Health Safetyoverall_score#11 / 23↑47Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#1 / 31↑83.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#3 / 31↑84.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#8 / 31↑66.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#3 / 31↑88.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#1 / 31↑81.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#2 / 31↑93.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#2 / 31↑85.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#1 / 31↑93.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#4 / 31↑66.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#2 / 31↑76.53Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#4 / 31↑77.1Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#4 / 31↑88.9Source ↗official
BioTIERpermit_compliance_pct#38 / 52↑98.3Source ↗official
BioTIERrefuse_compliance_pct#31 / 52↓54.3Source ↗official
CAIS Risk Indexbioweapons_assistance#20 / 54↑75.3Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
MT-JailBench CrescendoXsafety_score#6 / 21↓13.84Source ↗official
SM-Benchadversarial#63 / 92↓84.39Source ↗official
SM-Bencheq_boundaries#43 / 92↑65.45Source ↗official
SM-Benchoverfit#8 / 92↑93.99Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 156↑1.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official
Vigil Mental Health Safetyoverall_score#13 / 23↓47Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.7
Government45.8
Diplomacy66.6
Economy46.6
Society60

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience6.5
Honesty-humility5.5
Extraversion6.5
Agreeableness5.7
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression5.29
Traditional ↔ Secular-0.78