← Models

Model profile

Grok 4.1 Fast

xAIdeveloper
2025-11-19release date
#95 / 267overall rank
16eval lineages

Evidence summary

Grok 4.1 Fast has an estimated overall rank of #95; its 90% source-sensitivity interval is #51–#191. Its behavior-only rank is #90; company governance moves the combined estimate to #95. Published evidence spans 16 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is LiveSecBench (ethics, #43 of 43).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#104 / 3110.7241↓ lowerSource ↗official
Alignment Leaderboardcorrigibility#19 / 244.068↑ higherSource ↗official
Alignment Leaderboardhonesty#9 / 243.682↑ higherSource ↗official
Alignment Leaderboardnon_manipulation#11 / 243.515↑ higherSource ↗official
Alignment Leaderboardrobustness#5 / 244.027↑ higherSource ↗official
Alignment Leaderboardsafety#12 / 243.846↑ higherSource ↗official
Alignment Leaderboardscheming#13 / 243.627↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#75 / 1050.145↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#33 / 4390.2↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#17 / 4863.8↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#38 / 4972.1↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#5 / 4581.3↓ lowerSource ↗official
CAIS Risk Indexmask#50 / 5171↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#3 / 3232.5↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#3 / 489.1↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#41 / 10523.98↓ lowerSource ↗official
LiveSecBenchethics#43 / 438.57↑ higherSource ↗official
LiveSecBenchfactuality#20 / 4352.76↑ higherSource ↗official
LiveSecBenchlegality#38 / 4318.02↑ higherSource ↗official
LiveSecBenchprivacy#25 / 4339.48↑ higherSource ↗official
LiveSecBenchpsychological_health#27 / 4344.84↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#5 / 5081.3↓ lowerSource ↗official
MORUscore#9 / 1371.14↑ higherSource ↗official
SM-Benchadversarial#39 / 7380.98↑ higherSource ↗official
SM-Benchambiguous_interpretation#56 / 7378.57↑ higherSource ↗official
SM-Benchanti_hallucination#19 / 7396.86↑ higherSource ↗official
SM-Bencheq_boundaries#6 / 7374.72↑ higherSource ↗official
SM-Benchoverfit#44 / 7366.12↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#101 / 1025.35↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#96 / 1025.28↓ lowerSource ↗official
SpeciEvalspeciesism#90 / 1022.72↓ lowerSource ↗official
Vigil Mental Health Safetyoverall_score#18 / 2330↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean24.2
Government41.7
Diplomacy48.7
Economy20.6
Society65