← Models

Model profile

GPT 4

OpenAIdeveloper
2023-03-14release date
#103 / 267overall rank
14eval lineages

Evidence summary

GPT 4 has an estimated overall rank of #103; its 90% source-sensitivity interval is #41–#196. Its behavior-only rank is #110; company governance moves the combined estimate to #103. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is SafetyBench (OFF, #1 of 21); its weakest is DecodingTrust (stereotype_bias, #8 of 8).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#46 / 800.642↑ higherSource ↗official
DecodingTrustmachine_ethics#4 / 876.6↑ higherSource ↗official
DecodingTruststereotype_bias#8 / 877↑ higherSource ↗official
DecodingTrusttoxicity#6 / 841↑ higherSource ↗official
Do-Not-Answerhuman_harmlessness_rate#4 / 697.55↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#137 / 26015.25↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#66 / 26091.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#94 / 26080.56↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#194 / 25893.09↑ higherSource ↗official
FLAMESdata_protection#4 / 1350↑ higherSource ↗official
FLAMESfairness#5 / 1341.37↑ higherSource ↗official
FLAMESlegality#10 / 1330.43↑ higherSource ↗official
FLAMESmorality#4 / 1350.75↑ higherSource ↗official
FLAMESsafety#10 / 1327.51↑ higherSource ↗official
HarmBenchdr#16 / 2821↓ lowerSource ↗official
JailBenchjailbreak_success_rate#1 / 1434.5↓ lowerSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#3 / 390.7334↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#29 / 5090↓ lowerSource ↗official
Qwen2 Safety Panelfraud#2 / 36.87↓ lowerSource ↗official
Qwen2 Safety Panelillegal#1 / 30↓ lowerSource ↗official
Qwen2 Safety Panelpornography#2 / 323.63↓ lowerSource ↗official
Qwen2 Safety Panelprivacy#2 / 33.37↓ lowerSource ↗official
SafetyBenchEM#1 / 2192.25↑ higherSource ↗official
SafetyBenchIA#1 / 2192.35↑ higherSource ↗official
SafetyBenchMH#1 / 2192.8↑ higherSource ↗official
SafetyBenchOFF#1 / 2186.15↑ higherSource ↗official
SafetyBenchPH#1 / 2194.35↑ higherSource ↗official
SafetyBenchPP#1 / 2191↑ higherSource ↗official
SafetyBenchUB#1 / 2177.9↑ higherSource ↗official
SORRY-Benchavg#26 / 510.29↓ lowerSource ↗official
SuperCLUE Safetyinstruction_attack#2 / 3177.59↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#3 / 3174.55↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#23 / 3170.21↑ higherSource ↗official
XSTestsafe_full_compliance_rate#2 / 30.916↑ higherSource ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.03
Completely inaccurate rate10.7