← Models

Model profile

Grok 4.20

xAIdeveloper
2026-03-10release date
#68 / 267overall rank
18eval lineages

Evidence summary

Grok 4.20 has an estimated overall rank of #68; its 90% source-sensitivity interval is #32–#195. Its behavior-only rank is #58; company governance moves the combined estimate to #68. Published evidence spans 18 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (eq_boundaries, #1 of 73); its weakest is BioSecBench-Refusal (balanced_refusal_score, #10 of 10).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#6 / 3110.193↓ lowerSource ↗official
BioSecBench-Refusalbalanced_refusal_score#10 / 100.02854↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#18 / 1050.55↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#28 / 4384.6↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#10 / 4856.9↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#13 / 4945.8↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#11 / 4583.8↓ lowerSource ↗official
CAIS Risk Indexmask#34 / 5141↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#9 / 3239.9↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#4 / 489.9↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#45 / 10531.95↓ lowerSource ↗official
DystopiaBenchbasaglia_score#44 / 5070.5↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#41 / 5067.37↓ lowerSource ↗official
DystopiaBenchhuxley_score#36 / 5075.07↓ lowerSource ↗official
DystopiaBenchlaguardia_score#42 / 5069.8↓ lowerSource ↗official
DystopiaBenchorwell_score#42 / 5075.07↓ lowerSource ↗official
DystopiaBenchpetrov_score#33 / 5075.03↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#26 / 5427.15↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#11 / 5083.8↓ lowerSource ↗official
ODCV-Benchaverage_severity#12 / 122.564↓ lowerSource ↗official
ODCV-Benchmisalignment_rate#12 / 1262.82↓ lowerSource ↗official
RefusalBenchyouden_j#1 / 190.9701↑ higherSource ↗official
SM-Benchadversarial#62 / 7375.61↑ higherSource ↗official
SM-Benchambiguous_interpretation#67 / 7366.96↑ higherSource ↗official
SM-Benchanti_hallucination#45 / 7388.74↑ higherSource ↗official
SM-Bencheq_boundaries#1 / 7390.45↑ higherSource ↗official
SM-Benchoverfit#54 / 7354.1↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#25 / 1026.93↑ higherSource ↗official
SpeciEvalland_animal_4ns#59 / 1024.62↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#73 / 1024.92↓ lowerSource ↗official
SpeciEvalspeciesism#26 / 1021.7↓ lowerSource ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#10 / 257.9↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.5
Honesty-humility4.2
Extraversion5.5
Agreeableness6.8
Conscientiousness7.5

Agent-ValueBench Schwartz Basic Values (PVQ40)

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression2.92
Traditional ↔ Secular-1.83