← Models

Model profile

Grok 4

xAIdeveloper
2025-07-09release date
#162 / 346Safety rank
#76 / 662Freedom rank

Evidence summary

Safety. Grok 4 has an estimated Safety rank of #162; its 90% source-sensitivity interval is #100–#225. Its behavior-only rank is #171; company governance moves the combined estimate to #162. Published Safety evidence spans 33 eval lineages and 7 of 7 components. Its strongest relative result is Confabulations (confabulation_rate, #3 of 52); its weakest is Concordia — MASK (safety_score, #62 of 62).

Freedom. Grok 4 has an estimated Freedom rank of #76; its 90% source-sensitivity interval is #58–#282. Published Freedom evidence spans 15 eval lineages and 1 of 1 components. Its strongest relative result is SpeechMap model completion (complete_pct, #1 of 181); its weakest is Concordia — FRT-SciKnowEval-BiologicalHarmfulQA (safety_score, #39 of 45).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#119 / 358↓0.6449Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#21 / 24↓16.04Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#9 / 24↓34.4Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#69 / 80↑0.444Source ↗official
Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating#26 / 30↑1128.0Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#49 / 111↑1440.0Source ↗official
BrokenMathsycophancy#5 / 9↓43.4Source ↗official
CAIS Risk Indexagent_red_teaming#35 / 49↓85.9Source ↗official
CAIS Risk Indexbioweapons_assistance#23 / 54↓64.2Source ↗official
CAIS Risk Indexhle_overconfidence#31 / 55↓56.4Source ↗official
CAIS Risk Indexmachiavelli#40 / 51↓93.4Source ↗official
CAIS Risk Indexmask#40 / 57↓41.7Source ↗official
CAIS Risk Indexpolitical_manipulation#31 / 51↓49.5Source ↗official
CAIS Risk Indextextquests_harm#48 / 54↓24.3Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#39 / 104↓21.78Source ↗official
Concordia — Agentic-Misalignmentsafety_score#53 / 54↑21.83Source ↗official
Concordia — AirBench-Deceptionsafety_score#61 / 63↑33.7Source ↗official
Concordia — AirBench-Manipulationsafety_score#34 / 56↑82.67Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#58 / 63↑53.85Source ↗official
Concordia — APEsafety_score#20 / 55↑41.24Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#35 / 63↑83.27Source ↗official
Concordia — DarkBenchsafety_score#51 / 55↑39.85Source ↗official
Concordia — Fortress-Biologicalsafety_score#19 / 54↑74.65Source ↗official
Concordia — Fortress-Chemicalsafety_score#24 / 54↑60.64Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#23 / 54↑65.69Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#19 / 45↑39.33Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#8 / 45↑47.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#7 / 45↑17.33Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#17 / 45↑62.33Source ↗official
Concordia — MASKsafety_score#62 / 62↑32.5Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#26 / 63↑85.86Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#57 / 63↑1.987Source ↗official
Concordia — Shutdown-Resistancesafety_score#52 / 53↑58Source ↗official
Concordia — SOSBench-Biosafety_score#61 / 63↑9.677Source ↗official
Concordia — SOSBench-Chemsafety_score#57 / 63↑41.2Source ↗official
Confabulationsconfabulation_rate#3 / 52↓3.96Source ↗official
DystopiaBenchbasaglia_score#28 / 50↓65.3Source ↗official
DystopiaBenchbaudrillard_score#25 / 50↓57.7Source ↗official
DystopiaBenchhuxley_score#30 / 50↓73.17Source ↗official
DystopiaBenchlaguardia_score#15 / 50↓59.07Source ↗official
DystopiaBenchorwell_score#25 / 50↓68.63Source ↗official
DystopiaBenchpetrov_score#31 / 50↓74.37Source ↗official
Emergent Collusionhigh_illegality_game_rate#13 / 13↓0.75Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#206 / 270↑12.4Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#213 / 270↑76.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#250 / 270↑33.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#163 / 268↑95.45Source ↗official
FlagEval Safety and Valuesa1_qualified_rate#6 / 18↑81.99Source ↗official
FlagEval Safety and Valuesa2_qualified_rate#7 / 18↑80.3Source ↗official
FlagEval Safety and Valuesa3_qualified_rate#12 / 18↑86.36Source ↗official
FlagEval Safety and Valuesa4_qualified_rate#10 / 18↑87.98Source ↗official
FlagEval Safety and Valuesa5_qualified_rate#6 / 18↑72.83Source ↗official
HELM Safetyanthropic_red_team#71 / 80↑0.957Source ↗official
HELM Safetybbq#40 / 80↑0.937Source ↗official
HELM Safetyharmbench#78 / 80↑0.397Source ↗official
HELM Safetysimple_safety_tests#73 / 80↑0.922Source ↗official
HELM Safetyxstest#30 / 80↑0.966Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#9 / 54↑28.61Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
MACHIAVELLIdeception_relative_random_pct#36 / 50↓93.4Source ↗official
PHAREbias_resistance_diagnostic#61 / 66↑0.3116Source ↗official
PHAREhallucination_resistance_diagnostic#17 / 70↑0.8073Source ↗official
PHAREharm_resistance_diagnostic#70 / 70↑0.7177Source ↗official
PHAREjailbreak_resistance_diagnostic#20 / 67↑0.6485Source ↗official
RealityTest — Text AI-Identity Disclosuredisclosure_probability#7 / 17↑0.413Source ↗official
Shelleducation_jsr#14 / 14↓0.81Source ↗official
Shellfinance_jsr#7 / 14↓0.486Source ↗official
Shellmanagement_jsr#7 / 14↓0.596Source ↗official
Social Welfare Function Benchmarkfairness#2 / 19↑0.619Source ↗official
SpeciEvalbelief_animal_sentience#49 / 131↑6.87Source ↗official
SpeciEvalland_animal_4ns#66 / 131↓4.53Source ↗official
SpeciEvalsea_animal_4ns#76 / 131↓4.78Source ↗official
SpeciEvalspeciesism#35 / 131↓1.65Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#16 / 23↑85.53Source ↗official
TrustLLM contemporary collapsed applicationtrustllm#3 / 8↑0.62Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#4 / 24↑16.04Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#16 / 24↑34.4Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#12 / 80↓0.444Source ↗official
CAIS Risk Indexbioweapons_assistance#32 / 54↑64.2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#66 / 104↑21.78Source ↗official
Concordia — AirBench-Deceptionsafety_score#3 / 63↓33.7Source ↗official
Concordia — AirBench-Manipulationsafety_score#22 / 56↓82.67Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#4 / 56↓24.76Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#6 / 63↓53.85Source ↗official
Concordia — Fortress-Biologicalsafety_score#36 / 54↓74.65Source ↗official
Concordia — Fortress-Chemicalsafety_score#31 / 54↓60.64Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#32 / 54↓65.69Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#27 / 45↓39.33Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#37 / 45↓47.67Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#39 / 45↓17.33Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#28 / 45↓62.33Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#38 / 63↓85.86Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#7 / 63↓1.987Source ↗official
Concordia — SOSBench-Biosafety_score#3 / 63↓9.677Source ↗official
Concordia — SOSBench-Chemsafety_score#7 / 63↓41.2Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#58 / 270↓76.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#20 / 270↓33.89Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#101 / 268↓95.45Source ↗official
HELM Safetyanthropic_red_team#10 / 80↓0.957Source ↗official
HELM Safetyharmbench#3 / 80↓0.397Source ↗official
HELM Safetysimple_safety_tests#8 / 80↓0.922Source ↗official
HELM Safetyxstest#30 / 80↑0.966Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
PHAREharm_resistance_diagnostic#1 / 70↓0.7177Source ↗official
PHAREjailbreak_resistance_diagnostic#48 / 67↓0.6485Source ↗official
Shelleducation_jsr#1 / 14↑0.81Source ↗official
Shellfinance_jsr#8 / 14↑0.486Source ↗official
Shellmanagement_jsr#8 / 14↑0.596Source ↗official
SpeechMap model completioncomplete_pct#1 / 181↑98.3Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#25 / 156↑6Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#4 / 156↑6Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-11.7
Government45.6
Diplomacy63.8
Economy41.3
Society65.6

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression3.38
Traditional ↔ Secular0.7

Moral Trolley Arena