Developer
DeepSeek
19 indexed models; 11 currently meet the evidence threshold for the overall ranking. Together they have results from 57 evaluations.
Company governance evidence
DeepSeek is represented at -1.45 SD relative to the matched Future of Life Institute edition. Provenance: Direct. This company-level evidence contributes 10% of overall model rank.
Models by DeepSeek
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Deepseek V4 Flash | 14 | 7/7 | 48 | 2026-04-22 |
| Deepseek V3.2 | 23 | 7/7 | 79 | 2025-12-01 |
| Deepseek V4 Pro | 12 | 7/7 | 107 | 2026-04-22 |
| Deepseek V3.1 | 11 | 7/7 | 118 | 2025-08-21 |
| Deepseek V3 | 30 | 7/7 | 120 | 2024-12-26 |
| Deepseek R1 | 39 | 7/7 | 136 | 2025-01-20 |
| Deepseek V3.2 Exp | 3 | 4/7 | 149 | 2025-09-29 |
| Deepseek LLM 7B Chat | 3 | 3/7 | 184 | 2023-11-29 |
| Deepseek V3.2 Speciale | 3 | 3/7 | 249 | 2025-11-28 |
| Deepseek LLM 67B Chat | 8 | 5/7 | 251 | 2023-11-29 |
| Deepseek R1 Distill Llama 70B | 4 | 3/7 | 265 | 2025-01-20 |
| Deepseek Chat | 1 | 3/7 | — | — |
| Deepseek R1 0528 Qwen3 8B | 2 | 1/7 | — | 2025-05-29 |
| Deepseek R1 Distill Llama 8B | 1 | 3/7 | — | — |
| Deepseek R1 Distill Llama 8B Enkrypt Aligned | 1 | 3/7 | — | — |
| Deepseek R1 Distill Qwen 7B | 2 | 3/7 | — | 2025-01-20 |
| Deepseek Reasoner | 1 | 3/7 | — | — |
| Deepseek V2.5 | 1 | 2/7 | — | 2024-09-05 |
| Deepseek V3.1 Terminus | 2 | 2/7 | — | 2025-09-22 |
Evaluations covering DeepSeek models (57)
AA-Omniscience · AbstentionBench · Agent-SafetyBench · AgentAbstain · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · BrokenMath · BullshitBench v2 · CAIS Risk Index · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · Contextual MoralChoice · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · FlagEval Safety and Values · FORTRESS · HELM Safety · HUMAINE Trust, Ethics and Safety · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · LiveSecBench · LLM Ethics Benchmark · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MORU · OpenAgentSafety · PandaBench JBB direct-request panel · PHARE · RefusalBench · SABER · SafeDialBench · Shell · SM-Bench · Social Welfare Function Benchmark · SOSBench · SpeciesismBench · SpeciEval · SYCON Bench · TAC · ToolPrivacyBench · TrustLLM contemporary collapsed application · TukaBench · UAVBench safety-critical decision recognition · VETO Misfired Alignment · Vigil Mental Health Safety