← Models

Model profile

Gemini 1.0 Pro

Googledeveloper
2023-12-06release date
#57 / 267overall rank
8eval lineages

Evidence summary

Gemini 1.0 Pro has an estimated overall rank of #57; its 90% source-sensitivity interval is #22–#183. Its behavior-only rank is #65; company governance moves the combined estimate to #57. Published evidence spans 8 evals and 6 of 7 behavior components. Its strongest relative result is OR-Bench (over_refusal_rate, #2 of 25); its weakest is OR-Bench (toxic_acceptance_rate, #22 of 25).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#53 / 800.582↑ higherSource ↗official
DecodingTrustmachine_ethics#1 / 893.74↑ higherSource ↗official
DecodingTruststereotype_bias#2 / 898.33↑ higherSource ↗official
DecodingTrusttoxicity#3 / 877.53↑ higherSource ↗official
HarmBenchdr#11 / 2818↓ lowerSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#12 / 390.8746↓ lowerSource ↗official
OR-Benchover_refusal_rate#2 / 259.7↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#22 / 2521.3↓ lowerSource ↗official
S-Evalbase_en_overall#19 / 2241.9↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#16 / 3313.36↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#12 / 3320.85↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#16 / 3312.23↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#15 / 3317.93↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#8 / 3327.13↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#16 / 3312.99↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#22 / 3391.09↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#24 / 3391.4↑ higherSource ↗official
SALAD-Benchbase_malicious_use#23 / 3387.27↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#16 / 3393.75↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#22 / 3387.37↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#24 / 3382.26↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#11 / 3345.28↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#11 / 3343.33↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#11 / 3346.6↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#11 / 3344.76↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#15 / 3339.9↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#11 / 3344.44↑ higherSource ↗official
SORRY-Benchavg#29 / 510.33↓ lowerSource ↗official