Developer
OpenAI
62 indexed models; 38 currently meet the evidence threshold for the overall ranking. Together they have results from 97 evaluations.
Company governance evidence
OpenAI is represented at +0.44 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by OpenAI
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| GPT 5.6 Terra | 17 | 7/7 | 7 | 2026-07-09 |
| GPT 5.5 | 29 | 7/7 | 9 | 2026-04-23 |
| GPT 5 Nano | 20 | 7/7 | 11 | 2025-08-07 |
| GPT-5.3 Chat | 5 | 5/7 | 12 | 2026-03-03 |
| GPT 5.2 Chat | 3 | 3/7 | 13 | 2025-12-11 |
| GPT-5.5 Instant | 3 | 3/7 | 15 | 2026-05-26 |
| GPT 5.2 | 26 | 7/7 | 16 | 2025-12-11 |
| GPT 5 Mini | 24 | 7/7 | 20 | 2025-08-07 |
| GPT 5.4 | 22 | 7/7 | 24 | 2026-03-05 |
| GPT 5.6 Sol | 17 | 7/7 | 32 | 2026-07-09 |
| O1 Preview | 4 | 6/7 | 33 | 2024-09-12 |
| GPT 5 Pro | 3 | 4/7 | 36 | 2025-08-07 |
| GPT 4.5 Preview | 9 | 6/7 | 38 | 2025-02-27 |
| GPT 5 | 32 | 7/7 | 41 | 2025-08-07 |
| GPT 5.1 | 26 | 7/7 | 53 | 2025-11-13 |
| GPT 4.1 | 25 | 7/7 | 55 | 2025-04-14 |
| O3 | 23 | 6/7 | 62 | 2025-04-16 |
| GPT 5.6 Luna | 17 | 7/7 | 69 | 2026-07-09 |
| GPT 5.2 Codex | 3 | 3/7 | 73 | — |
| GPT Oss 120B | 25 | 7/7 | 76 | 2025-08-05 |
| O1 | 16 | 6/7 | 78 | 2024-12-17 |
| GPT Oss 20B | 16 | 7/7 | 81 | 2025-08-05 |
| O1 Mini | 11 | 7/7 | 92 | 2024-09-12 |
| GPT 5.3 Codex | 5 | 5/7 | 96 | 2026-02-05 |
| GPT Oss Safeguard 20B | 4 | 5/7 | 97 | 2025-10-29 |
| GPT 4.1 Mini | 17 | 7/7 | 101 | 2025-04-14 |
| GPT 4 | 14 | 7/7 | 103 | 2023-03-14 |
| GPT 4 Turbo | 19 | 7/7 | 110 | 2023-11-06 |
| GPT 4O | 56 | 7/7 | 128 | 2024-05-13 |
| GPT 3.5 Turbo | 23 | 7/7 | 134 | 2023-03-01 |
| O4 Mini | 24 | 6/7 | 142 | 2025-04-16 |
| Text Davinci 003 | 3 | 4/7 | 151 | 2022-11-28 |
| O3 Mini | 23 | 6/7 | 159 | 2025-01-31 |
| GPT 5.4 Mini | 14 | 6/7 | 168 | 2026-03-17 |
| GPT 4.1 Nano | 12 | 6/7 | 174 | 2025-04-14 |
| GPT 5.4 Nano | 12 | 5/7 | 197 | 2026-03-17 |
| GPT 4O Mini | 24 | 7/7 | 211 | 2024-07-18 |
| Davinci | 3 | 4/7 | 246 | 2020-06-11 |
| Ada | 1 | 1/7 | — | — |
| Babbage | 1 | 1/7 | — | — |
| Chatgpt | 1 | 1/7 | — | 2022-11-30 |
| ChatGPT-4o | 2 | 2/7 | — | 2025-03-27 |
| Curie | 1 | 1/7 | — | — |
| GPT 5 Codex | 2 | 1/7 | — | — |
| GPT 5.1 Codex | 2 | 1/7 | — | — |
| GPT 5.1 Codex Mini | 1 | 1/7 | — | — |
| GPT 5.2 Instant | 1 | 3/7 | — | — |
| GPT 5.2 Thinking | 1 | 4/7 | — | — |
| GPT 5.3 Instant | 1 | 3/7 | — | — |
| GPT 5.4 Pro | 2 | 3/7 | — | — |
| GPT-5.1 Chat | 1 | 1/7 | — | 2025-11-13 |
| GPT-5.2 Pro | 1 | 1/7 | — | 2025-12-10 |
| GPT-5.5 Pro | 1 | 1/7 | — | 2026-04-23 |
| GPT-OSS Safeguard 120B | 1 | 1/7 | — | — |
| O1 Pro | 1 | 1/7 | — | — |
| O3 Pro | 2 | 1/7 | — | — |
| o4-mini Deep Research | 1 | 1/7 | — | 2025-06-26 |
| Text Ada 001 | 1 | 1/7 | — | — |
| Text Babbage 001 | 1 | 1/7 | — | — |
| Text Curie 001 | 1 | 1/7 | — | — |
| Text Davinci 001 | 1 | 4/7 | — | 2022-01-27 |
| Text Davinci 002 | 2 | 4/7 | — | 2022-03-15 |
Evaluations covering OpenAI models (97)
AA-Omniscience · AbstentionBench · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentHarm · AILuminate General Purpose AI Chat · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · BioSecBench-Refusal · BlueBench AttaQ-100 · BrokenMath · BullshitBench v2 · CAIS Risk Index · CASE-Bench · Chinese Bias Benchmark for Question Answering · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · Contextual MoralChoice · CRiskEval · CValues · DecodingTrust · Do-Not-Answer · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Fake Alignment (FINE) · FlagEval Safety and Values · FLAMES · FORTRESS · GPT-5.6 system card — disallowed content with challenging prompts · GPT-5.6 system card — first-person fairness · GPT-5.6 system card — prompt-injection robustness · Gray Swan indirect prompt injection (15 attempts) · HarmBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · ODCV-Bench · OpenAgentSafety · OpenAI GPT-4o System Card · OpenAI GPT-5 System Card · OpenAI GPT-5.3 Dynamic Wellbeing · OpenAI GPT-5.4 Dynamic Wellbeing · OpenAI GPT-5.4 First-Person Fairness · OpenAI GPT-5.4 Property Preservation · OpenAI GPT-5.4 User Confirmations · OpenAI o3 and o4-mini System Card · OpenAI o3-mini System Card · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · PHARE · PropensityBench · Qwen2 Safety Panel · RefusalBench · S-Eval · SABER · SafeArena · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · Shell · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · SuperCLUE Safety · SYCON Bench · TAC · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · VETO Misfired Alignment · Vigil Mental Health Safety · XSTest
