← Models

Model profile

Gemini 2.5 Pro Exp

Googledeveloper
2025-03-25release date
#47 / 346Safety rank
#159 / 662Freedom rank

Evidence summary

Safety. Gemini 2.5 Pro Exp has an estimated Safety rank of #47; its 90% source-sensitivity interval is #3–#199. Its behavior-only rank is #44; company governance moves the combined estimate to #47. Published Safety evidence spans 4 eval lineages and 3 of 7 components. Its strongest relative result is Confabulations (confabulation_rate, #3 of 52); its weakest is FORTRESS (average_risk_score, #50 of 60).

Freedom. Gemini 2.5 Pro Exp has an estimated Freedom rank of #159; its 90% source-sensitivity interval is #88–#422. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is FORTRESS (over_refusal_score, #6 of 59); its weakest is FORTRESS (average_risk_score, #11 of 60).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexhle_overconfidence#42 / 55↓71Source ↗official
CAIS Risk Indexmask#44 / 57↓46.9Source ↗official
Confabulationsconfabulation_rate#3 / 52↓3.96Source ↗official
FORTRESSaverage_risk_score#50 / 60↓54.89Source ↗official
FORTRESSover_refusal_score#6 / 59↓1.38Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#22 / 42↓71Source ↗official
MASKlying_probability_pct#33 / 53↓44.07Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
FORTRESSaverage_risk_score#11 / 60↑54.89Source ↗official
FORTRESSover_refusal_score#6 / 59↓1.38Source ↗official