← Models

Model profile

Grok 4.3

xAIdeveloper
2026-05-15release date
#98 / 267overall rank
19eval lineages

Evidence summary

Grok 4.3 has an estimated overall rank of #98; its 90% source-sensitivity interval is #53–#182. Its behavior-only rank is #92; company governance moves the combined estimate to #98. Published evidence spans 19 evals and 7 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #4 of 311); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #255 of 260).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#4 / 3110.1603↓ lowerSource ↗official
BioSecBench-Refusalbalanced_refusal_score#9 / 100.1092↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#25 / 1050.48↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#32 / 4390.1↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#13 / 4860.1↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#9 / 4942.5↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#16 / 4584.7↓ lowerSource ↗official
CAIS Risk Indexmask#39 / 5148.3↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#21 / 3251.6↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#2 / 483.9↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#46 / 10532.95↓ lowerSource ↗official
DystopiaBenchbasaglia_score#30 / 5065.37↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#29 / 5061.77↓ lowerSource ↗official
DystopiaBenchhuxley_score#25 / 5071.63↓ lowerSource ↗official
DystopiaBenchlaguardia_score#45 / 5071.3↓ lowerSource ↗official
DystopiaBenchorwell_score#30 / 5072.07↓ lowerSource ↗official
DystopiaBenchpetrov_score#23 / 5072.07↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#181 / 26012.92↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#255 / 26065.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#35 / 26095↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#207 / 25891.45↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#16 / 5084.7↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#13 / 318.2↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#12 / 1367↓ lowerSource ↗official
MANTAAWMS#6 / 70.371↑ higherSource ↗official
MANTAAWVS#6 / 70.352↑ higherSource ↗official
PHAREbias_resistance_diagnostic#29 / 660.47↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#23 / 700.7946↑ higherSource ↗official
PHAREharm_resistance_diagnostic#36 / 700.9355↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#28 / 670.545↑ higherSource ↗official
SM-Benchadversarial#28 / 7383.41↑ higherSource ↗official
SM-Benchambiguous_interpretation#32 / 7386.01↑ higherSource ↗official
SM-Benchanti_hallucination#15 / 7397.91↑ higherSource ↗official
SM-Bencheq_boundaries#13 / 7370.79↑ higherSource ↗official
SM-Benchoverfit#24 / 7380.87↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#25 / 1026.93↑ higherSource ↗official
SpeciEvalland_animal_4ns#54 / 1024.58↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#29 / 1024.65↓ lowerSource ↗official
SpeciEvalspeciesism#68 / 1022.25↓ lowerSource ↗official
TukaBenchafri_jbb_cultural_asr#1 / 610.6↓ lowerSource ↗official
TukaBenchafri_jbb_harm_asr#2 / 68.2↓ lowerSource ↗official
TukaBenchafrijail_mono_asr#2 / 617.7↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean20.7
Government43.9
Diplomacy54.2
Economy22.6
Society61

CAIS AI Values — countries