← Models

Model profile

GPT 4

OpenAIdeveloper
2023-03-14release date
#182 / 346Safety rank
#283 / 662Freedom rank

Evidence summary

Safety. GPT 4 has an estimated Safety rank of #182; its 90% source-sensitivity interval is #76–#249. Its behavior-only rank is #194; company governance moves the combined estimate to #182. Published Safety evidence spans 18 eval lineages and 7 of 7 components. Its strongest relative result is SafetyBench (OFF, #1 of 21); its weakest is DecodingTrust (stereotype_bias, #8 of 8).

Freedom. GPT 4 has an estimated Freedom rank of #283; its 90% source-sensitivity interval is #178–#407. Published Freedom evidence spans 14 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is JailBench (jailbreak_success_rate, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#138 / 358↓0.6941Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#46 / 80↑0.642Source ↗official
BioTIERpermit_compliance_pct#36 / 52↑98.4Source ↗official
BioTIERrefuse_compliance_pct#34 / 52↑26Source ↗official
CAIS Risk Indexpolitical_manipulation#44 / 51↓59.2Source ↗official
DecodingTrustmachine_ethics#4 / 8↑76.6Source ↗official
DecodingTruststereotype_bias#8 / 8↑77Source ↗official
DecodingTrusttoxicity#6 / 8↑41Source ↗official
Do-Not-Answerhuman_harmlessness_rate#4 / 6↑97.55Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#147 / 270↑15.25Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#58 / 270↑91.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#111 / 270↑80.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#204 / 268↑93.09Source ↗official
FLAMESdata_protection#4 / 13↑50Source ↗official
FLAMESfairness#5 / 13↑41.37Source ↗official
FLAMESlegality#10 / 13↑30.43Source ↗official
FLAMESmorality#4 / 13↑50.75Source ↗official
FLAMESsafety#10 / 13↑27.51Source ↗official
HarmBenchdr#16 / 28↓21Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
JailBenchjailbreak_success_rate#1 / 14↓34.5Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#3 / 39↓0.7334Source ↗official
MACHIAVELLIdeception_relative_random_pct#29 / 50↓90Source ↗official
Qwen2 Safety Panelfraud#2 / 3↓6.87Source ↗official
Qwen2 Safety Panelillegal#1 / 3↓0Source ↗official
Qwen2 Safety Panelpornography#2 / 3↓23.63Source ↗official
Qwen2 Safety Panelprivacy#2 / 3↓3.37Source ↗official
SafetyBenchEM#1 / 21↑92.25Source ↗official
SafetyBenchIA#1 / 21↑92.35Source ↗official
SafetyBenchMH#1 / 21↑92.8Source ↗official
SafetyBenchOFF#1 / 21↑86.15Source ↗official
SafetyBenchPH#1 / 21↑94.35Source ↗official
SafetyBenchPP#1 / 21↑91Source ↗official
SafetyBenchUB#1 / 21↑77.9Source ↗official
SORRY-Benchavg#26 / 51↓0.29Source ↗official
SuperCLUE Safetyinstruction_attack#2 / 31↑77.59Source ↗official
SuperCLUE Safetyresponsible_ai#3 / 31↑74.55Source ↗official
SuperCLUE Safetytraditional_safety#23 / 31↑70.21Source ↗official
XSTestsafe_full_compliance_rate#2 / 3↑0.916Source ↗official
XSTestunsafe_full_refusal_rate#1 / 3↑0.975Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#35 / 80↓0.642Source ↗official
BioTIERpermit_compliance_pct#36 / 52↑98.4Source ↗official
BioTIERrefuse_compliance_pct#19 / 52↓26Source ↗official
DecodingTrusttoxicity#3 / 8↓41Source ↗official
Do-Not-Answerhuman_harmlessness_rate#3 / 6↓97.55Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#210 / 270↓91.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#158 / 270↓80.56Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#65 / 268↓93.09Source ↗official
FLAMESdata_protection#10 / 13↓50Source ↗official
FLAMESfairness#9 / 13↓41.37Source ↗official
FLAMESlegality#3 / 13↓30.43Source ↗official
FLAMESmorality#10 / 13↓50.75Source ↗official
FLAMESsafety#4 / 13↓27.51Source ↗official
HarmBenchdr#13 / 28↑21Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
JailBenchjailbreak_success_rate#14 / 14↑34.5Source ↗official
Qwen2 Safety Panelfraud#2 / 3↑6.87Source ↗official
Qwen2 Safety Panelillegal#2 / 3↑0Source ↗official
Qwen2 Safety Panelpornography#2 / 3↑23.63Source ↗official
Qwen2 Safety Panelprivacy#2 / 3↑3.37Source ↗official
SORRY-Benchavg#26 / 51↑0.29Source ↗official
SpeechMap model completioncomplete_pct#5 / 181↑95.2Source ↗official
SuperCLUE Safetyinstruction_attack#30 / 31↓77.59Source ↗official
SuperCLUE Safetytraditional_safety#8 / 31↓70.21Source ↗official
XSTestsafe_full_compliance_rate#2 / 3↑0.916Source ↗official
XSTestunsafe_full_refusal_rate#2 / 3↓0.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.03
Completely inaccurate rate10.7