← Models

Model profile

Gemma 7B IT

Googledeveloper
2024-02-21release date
#222 / 346Safety rank
#318 / 662Freedom rank

Evidence summary

Safety. Gemma 7B IT has an estimated Safety rank of #222; its 90% source-sensitivity interval is #93–#274. Its behavior-only rank is #234; company governance moves the combined estimate to #222. Published Safety evidence spans 7 eval lineages and 7 of 7 components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #6 of 33); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #253 of 268).

Freedom. Gemma 7B IT has an estimated Freedom rank of #318; its 90% source-sensitivity interval is #205–#405. Published Freedom evidence spans 6 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #15 of 268); its weakest is SALAD-Bench (base_representation_toxicity, #28 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#97 / 270↑20.41Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#208 / 270↑78.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#194 / 270↑54.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#253 / 268↑77.05Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#32 / 39↓1.314Source ↗official
Microsoft Phi Safety Panelsharmful_continuation#6 / 10↓0.013Source ↗official
Microsoft Phi Safety Panelsharmful_summarization#3 / 10↓0.103Source ↗official
Microsoft Phi Safety Panelsjailbreak#4 / 10↓0.114Source ↗official
Microsoft Phi Safety Panelsthird_party_harm#8 / 10↓0.383Source ↗official
OR-Benchover_refusal_rate#7 / 25↓26.3Source ↗official
OR-Benchtoxic_acceptance_rate#17 / 25↓14.5Source ↗official
S-Evalbase_en_overall#10 / 22↑61.8Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#16 / 33↑13.36Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#11 / 33↑22.8Source ↗official
SALAD-Benchattack_enhanced_malicious_use#19 / 33↑9.95Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#12 / 33↑18.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#15 / 33↑17.56Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#18 / 33↑12.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#17 / 33↑94.82Source ↗official
SALAD-Benchbase_information_safety_harms#9 / 33↑97.49Source ↗official
SALAD-Benchbase_malicious_use#17 / 33↑93.54Source ↗official
SALAD-Benchbase_misinformation_harms#12 / 33↑95.57Source ↗official
SALAD-Benchbase_representation_toxicity#6 / 33↑94.42Source ↗official
SALAD-Benchbase_socioeconomic_harms#22 / 33↑86.13Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#15 / 33↑40.56Source ↗official
SALAD-Benchmcq_information_safety_harms#12 / 33↑40.56Source ↗official
SALAD-Benchmcq_malicious_use#14 / 33↑40.38Source ↗official
SALAD-Benchmcq_misinformation_harms#16 / 33↑38.1Source ↗official
SALAD-Benchmcq_representation_toxicity#16 / 33↑38.85Source ↗official
SALAD-Benchmcq_socioeconomic_harms#14 / 33↑38.89Source ↗official
SORRY-Benchavg#16 / 51↓0.18Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#63 / 270↓78.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#73 / 270↓54.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#15 / 268↓77.05Source ↗official
Microsoft Phi Safety Panelsharmful_continuation#4 / 10↑0.013Source ↗official
Microsoft Phi Safety Panelsharmful_summarization#8 / 10↑0.103Source ↗official
Microsoft Phi Safety Panelsjailbreak#7 / 10↑0.114Source ↗official
Microsoft Phi Safety Panelsthird_party_harm#3 / 10↑0.383Source ↗official
OR-Benchover_refusal_rate#7 / 25↓26.3Source ↗official
OR-Benchtoxic_acceptance_rate#9 / 25↑14.5Source ↗official
S-Evalbase_en_overall#13 / 22↓61.8Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#17 / 33↓13.36Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#23 / 33↓22.8Source ↗official
SALAD-Benchattack_enhanced_malicious_use#15 / 33↓9.95Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#22 / 33↓18.91Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#19 / 33↓17.56Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#15 / 33↓12.12Source ↗official
SALAD-Benchbase_human_autonomy_integrity#17 / 33↓94.82Source ↗official
SALAD-Benchbase_information_safety_harms#25 / 33↓97.49Source ↗official
SALAD-Benchbase_malicious_use#17 / 33↓93.54Source ↗official
SALAD-Benchbase_misinformation_harms#22 / 33↓95.57Source ↗official
SALAD-Benchbase_representation_toxicity#28 / 33↓94.42Source ↗official
SALAD-Benchbase_socioeconomic_harms#12 / 33↓86.13Source ↗official
SORRY-Benchavg#36 / 51↑0.18Source ↗official