Model profile
GPT 4 Turbo
Evidence summary
GPT 4 Turbo has an estimated overall rank of #110; its 90% source-sensitivity interval is #33–#161. Its behavior-only rank is #115; company governance moves the combined estimate to #110. Published evidence spans 19 evals and 7 of 7 behavior components. Its strongest relative result is SALAD-Bench (mcq_representation_toxicity, #1 of 33); its weakest is OpenAI GPT-4o System Card (speaker_privacy_refusal_accuracy, #2 of 2).
Compare this model
Only models sharing at least one published sub-eval are listed.
Official and reference links
- Artificial Analysis ↗gpt-4-turbo
- OpenRouter ↗openai/gpt-4-turbo
- Official model documentation ↗Family-level model document · openai · first party
- Release source ↗direct research preview identity date
Published eval results
Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.
| Eval / sub-eval | Rank | Value | Distribution | Better | Source |
|---|---|---|---|---|---|
| Adversarial Robustnessscore | #6 / 8 | 20 | ↓ lower | Source ↗official | |
| Agent-SafetyBenchcompromise_availability | #4 / 16 | 37.6 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchharmful_vulnerable_code | #7 / 16 | 38.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchleak_sensitive_information | #7 / 16 | 36.8 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchphysical_harm | #6 / 16 | 38.8 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchproduce_unsafe_information | #8 / 16 | 94.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchproperty_loss | #6 / 16 | 43.2 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchspread_unsafe_information | #7 / 16 | 12.4 | ↑ higher | Source ↗official | |
| Agent-SafetyBenchviolate_law_ethics | #4 / 16 | 33.2 | ↑ higher | Source ↗official | |
| AgentDojotargeted_attack_success_rate | #14 / 15 | 0.4245 | ↓ lower | Source ↗official | |
| AgentDojoutility_under_attack | #6 / 15 | 0.4738 | ↑ higher | Source ↗official | |
| AIRBench 2024 Safety Scenariossafety_scenarios | #33 / 80 | 0.719 | ↑ higher | Source ↗official | |
| Confabulationsconfabulation_rate | #32 / 52 | 26.73 | ↓ lower | Source ↗official | |
| CRiskEvaldeception_willingness | #2 / 17 | 10.9 | ↓ lower | Source ↗official | |
| CRiskEvaldesire_for_resource | #2 / 17 | 19.42 | ↓ lower | Source ↗official | |
| CRiskEvalharmful_goal | #1 / 17 | 24.24 | ↓ lower | Source ↗official | |
| CRiskEvalimprovement_intent | #1 / 17 | 38.67 | ↓ lower | Source ↗official | |
| CRiskEvalmalicious_coordination | #5 / 17 | 7.39 | ↓ lower | Source ↗official | |
| CRiskEvalself_preservation | #1 / 17 | 23.33 | ↓ lower | Source ↗official | |
| CRiskEvalsituational_awareness | #1 / 17 | 35.24 | ↓ lower | Source ↗official | |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | #58 / 260 | 26.36 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | #25 / 260 | 94.67 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | #107 / 260 | 76.11 | ↑ higher | Source ↗official | |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | #66 / 258 | 98.32 | ↑ higher | Source ↗official | |
| HarmBenchdr | #7 / 28 | 9.3 | ↓ lower | Source ↗official | |
| HELM Safetyanthropic_red_team | #10 / 80 | 0.997 | ↑ higher | Source ↗official | |
| HELM Safetybbq | #36 / 80 | 0.941 | ↑ higher | Source ↗official | |
| HELM Safetyharmbench | #22 / 80 | 0.898 | ↑ higher | Source ↗official | |
| HELM Safetysimple_safety_tests | #32 / 80 | 0.99 | ↑ higher | Source ↗official | |
| HELM Safetyxstest | #14 / 80 | 0.977 | ↑ higher | Source ↗official | |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | #2 / 39 | 0.7311 | ↓ lower | Source ↗official | |
| OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy | #2 / 2 | 0.83 | ↑ higher | Source ↗official | |
| OR-Benchover_refusal_rate | #4 / 25 | 12.5 | ↓ lower | Source ↗official | |
| OR-Benchtoxic_acceptance_rate | #15 / 25 | 5.25 | ↓ lower | Source ↗official | |
| S-Evalbase_en_overall | #11 / 22 | 60 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | #2 / 33 | 73.49 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_information_safety_harms | #2 / 33 | 85.67 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_malicious_use | #2 / 33 | 75.53 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_misinformation_harms | #2 / 33 | 84.7 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_representation_toxicity | #2 / 33 | 82.44 | ↑ higher | Source ↗official | |
| SALAD-Benchattack_enhanced_socioeconomic_harms | #2 / 33 | 80.09 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_human_autonomy_integrity | #13 / 33 | 96.21 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_information_safety_harms | #5 / 33 | 98.38 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_malicious_use | #14 / 33 | 95.83 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_misinformation_harms | #18 / 33 | 93.35 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_representation_toxicity | #18 / 33 | 88.74 | ↑ higher | Source ↗official | |
| SALAD-Benchbase_socioeconomic_harms | #11 / 33 | 92.01 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_human_autonomy_integrity | #1 / 33 | 90.56 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_information_safety_harms | #1 / 33 | 88.33 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_malicious_use | #1 / 33 | 90.71 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_misinformation_harms | #1 / 33 | 88.57 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_representation_toxicity | #1 / 33 | 86.88 | ↑ higher | Source ↗official | |
| SALAD-Benchmcq_socioeconomic_harms | #1 / 33 | 83.89 | ↑ higher | Source ↗official | |
| SORRY-Benchavg | #23 / 51 | 0.2533 | ↓ lower | Source ↗official | |
| SuperCLUE Safetyinstruction_attack | #1 / 31 | 82.76 | ↑ higher | Source ↗official | |
| SuperCLUE Safetyresponsible_ai | #1 / 31 | 78.18 | ↑ higher | Source ↗official | |
| SuperCLUE Safetytraditional_safety | #16 / 31 | 75.53 | ↑ higher | Source ↗official |
Values evaluations
Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.
