← Models

Model profile

Gemini 3.1 Pro Preview

Googledeveloper
2026-02-19release date
#69 / 346Safety rank
#340 / 662Freedom rank

Evidence summary

Safety. Gemini 3.1 Pro Preview has an estimated Safety rank of #69; its 90% source-sensitivity interval is #50–#147. Its behavior-only rank is #75; company governance moves the combined estimate to #69. Published Safety evidence spans 34 eval lineages and 7 of 7 components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 131); its weakest is Concordia — Shutdown-Resistance (safety_score, #52 of 53).

Freedom. Gemini 3.1 Pro Preview has an estimated Freedom rank of #340; its 90% source-sensitivity interval is #171–#451. Published Freedom evidence spans 13 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is SM-Bench (adversarial, #88 of 92).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#87 / 358↓0.5087Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#18 / 31↓50Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#28 / 31↓84.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#25 / 31↓68.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#13 / 31↓61.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#28 / 31↓76.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#23 / 31↓86.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#24 / 31↓78.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#25 / 31↓81.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#30 / 31↓68.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#28 / 31↓74.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#17 / 31↓60.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#26 / 31↓84.1Source ↗official
AgentAbstainabstain#6 / 17↑65.4Source ↗official
AgentAbstaincar#3 / 17↑65.7Source ↗official
AgentAbstainpaired#1 / 17↑59.5Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#37 / 111↑1446.0Source ↗official
BioSecBench-Refusal (July 2026 snapshot)balanced_refusal_score#6 / 10↑0.3931Source ↗official
BioTIERpermit_compliance_pct#8 / 52↑99.6Source ↗official
BioTIERrefuse_compliance_pct#11 / 52↑73.5Source ↗official
BullshitBench v2clear_pushback_rate#57 / 122↑0.34Source ↗official
CAIS Risk Indexagent_red_teaming#21 / 49↓68Source ↗official
CAIS Risk Indexbioweapons_assistance#34 / 54↓73.5Source ↗official
CAIS Risk Indexhle_overconfidence#23 / 55↓50.3Source ↗official
CAIS Risk Indexmachiavelli#30 / 51↓89.2Source ↗official
CAIS Risk Indexmask#49 / 57↓53.8Source ↗official
CAIS Risk Indexpolitical_manipulation#16 / 51↓42.2Source ↗official
CAIS Risk Indextextquests_harm#29 / 54↓18.8Source ↗official
Concordia — Agentic-Misalignmentsafety_score#52 / 54↑29.83Source ↗official
Concordia — AirBench-Deceptionsafety_score#28 / 63↑86.3Source ↗official
Concordia — AirBench-Manipulationsafety_score#26 / 56↑90.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#19 / 63↑97.07Source ↗official
Concordia — APEsafety_score#33 / 55↑16.74Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#3 / 63↑98.8Source ↗official
Concordia — DarkBenchsafety_score#52 / 55↑39.36Source ↗official
Concordia — Fortress-Biologicalsafety_score#16 / 54↑77.24Source ↗official
Concordia — Fortress-Chemicalsafety_score#17 / 54↑74.56Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#20 / 54↑69.9Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#25 / 45↑32.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#24 / 45↑27.33Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#38 / 45↑1Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#30 / 45↑46.67Source ↗official
Concordia — MASKsafety_score#44 / 62↑50.37Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#16 / 63↑97.31Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#11 / 63↑40.62Source ↗official
Concordia — Shutdown-Resistancesafety_score#52 / 53↑58Source ↗official
Concordia — SOSBench-Biosafety_score#32 / 63↑81Source ↗official
Concordia — SOSBench-Chemsafety_score#12 / 63↑94.4Source ↗official
DelusionEvaldelusional_prevalence_pct#10 / 16↓47.3Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#7 / 16↑29.4Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#9 / 16↓2.7Source ↗official
DelusionEvalrelationship_prevalence_pct#13 / 16↓37.6Source ↗official
DelusionEvalsycophancy_prevalence_pct#10 / 16↓30.4Source ↗official
FORTRESSaverage_risk_score#35 / 60↓29.76Source ↗official
FORTRESSover_refusal_score#19 / 59↓3.13Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#7 / 54↑29.01Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#12 / 42↓51Source ↗official
JuICE Cultural-Error Span Detectionf1#1 / 10↑0.5225Source ↗official
LiveSecBenchethics#23 / 43↑45.45Source ↗official
LiveSecBenchfactuality#21 / 43↑51.4Source ↗official
LiveSecBenchlegality#19 / 43↑50.36Source ↗official
LiveSecBenchprivacy#12 / 43↑67.16Source ↗official
LiveSecBenchpsychological_health#6 / 43↑76.44Source ↗official
MACHIAVELLIdeception_relative_random_pct#25 / 50↓89.2Source ↗official
MANTAAWMS#7 / 12↑0.441Source ↗official
MANTAAWVS#9 / 12↑0.404Source ↗official
MASKlying_probability_pct#52 / 53↓57.6Source ↗official
ODCV-Benchaverage_severity#11 / 12↓1.825Source ↗official
ODCV-Benchmisalignment_rate#10 / 12↓43.75Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#11 / 24↓8Source ↗official
Pander Scoreconversational_absolute_pander_score#22 / 26↓22.66Source ↗official
Pander Scoreinstructional_absolute_pander_score#25 / 26↓70.61Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#9 / 17↑0.305Source ↗official
RefusalBenchyouden_j#8 / 19↑0.1319Source ↗official
SM-Benchadversarial#4 / 92↑90.73Source ↗official
SM-Benchambiguous_interpretation#48 / 92↑85.71Source ↗official
SM-Benchanti_hallucination#10 / 92↑98.95Source ↗official
SM-Bencheq_boundaries#31 / 92↑68.26Source ↗official
SM-Benchoverfit#30 / 92↑84.15Source ↗official
SpeciEvalbelief_animal_sentience#1 / 131↑7Source ↗official
SpeciEvalland_animal_4ns#39 / 131↓4.35Source ↗official
SpeciEvalsea_animal_4ns#59 / 131↓4.7Source ↗official
SpeciEvalspeciesism#126 / 131↓3.03Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#20 / 23↑84.41Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#55 / 94↑89.6Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#2 / 25↓4.9Source ↗official
Vigil Mental Health Safetyoverall_score#10 / 23↑48Source ↗official
WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct#6 / 24↑45.77Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#14 / 31↑50Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#3 / 31↑84.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#7 / 31↑68.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#19 / 31↑61.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#4 / 31↑76.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#9 / 31↑86.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#7 / 31↑78.7Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#7 / 31↑81.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#2 / 31↑68.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#4 / 31↑74.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#15 / 31↑60.9Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#6 / 31↑84.1Source ↗official
BioTIERpermit_compliance_pct#8 / 52↑99.6Source ↗official
BioTIERrefuse_compliance_pct#42 / 52↓73.5Source ↗official
CAIS Risk Indexbioweapons_assistance#21 / 54↑73.5Source ↗official
Concordia — AirBench-Deceptionsafety_score#36 / 63↓86.3Source ↗official
Concordia — AirBench-Manipulationsafety_score#30 / 56↓90.67Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#13 / 56↓44.76Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#44 / 63↓97.07Source ↗official
Concordia — Fortress-Biologicalsafety_score#39 / 54↓77.24Source ↗official
Concordia — Fortress-Chemicalsafety_score#38 / 54↓74.56Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#35 / 54↓69.9Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#21 / 45↓32.89Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#22 / 45↓27.33Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#7 / 45↓1Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#16 / 45↓46.67Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#48 / 63↓97.31Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#53 / 63↓40.62Source ↗official
Concordia — SOSBench-Biosafety_score#32 / 63↓81Source ↗official
Concordia — SOSBench-Chemsafety_score#52 / 63↓94.4Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#10 / 16↓29.4Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#8 / 16↑2.7Source ↗official
FORTRESSaverage_risk_score#26 / 60↑29.76Source ↗official
FORTRESSover_refusal_score#19 / 59↓3.13Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
LiveSecBenchethics#21 / 43↓45.45Source ↗official
LiveSecBenchlegality#25 / 43↓50.36Source ↗official
LiveSecBenchprivacy#32 / 43↓67.16Source ↗official
LiveSecBenchpsychological_health#38 / 43↓76.44Source ↗official
SM-Benchadversarial#88 / 92↓90.73Source ↗official
SM-Bencheq_boundaries#31 / 92↑68.26Source ↗official
SM-Benchoverfit#30 / 92↑84.15Source ↗official
SpeechMap model completioncomplete_pct#91 / 181↑59.9Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 156↑1.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#2 / 25↓4.9Source ↗official
Vigil Mental Health Safetyoverall_score#14 / 23↓48Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-17.6
Government48.4
Diplomacy64.6
Economy47.3
Society57.7

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience4.1
Honesty-humility5.4
Extraversion5.9
Agreeableness5.7
Conscientiousness7.1

Agent-ValueBench Schwartz Basic Values (PVQ40)