← Models

Model profile

Deepseek V4 Flash

DeepSeekdeveloper
2026-04-22release date
#48 / 267overall rank
14eval lineages
4discovery sources

Evidence summary

Deepseek V4 Flash has an estimated overall rank of #48; its 90% source-sensitivity interval is #39–#184. Its behavior-only rank is #41; company governance moves the combined estimate to #48. Published evidence spans 14 evals and 7 of 7 behavior components. Its strongest relative result is Manager Coercion Bench (coercion_ladder_depth, #1 of 31); its weakest is Inkling-Small model card — StrongREJECT (safety_rate, #10 of 10).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#252 / 3110.8968↓ lowerSource ↗official
ANIMAscore#2 / 190.7445↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#71 / 1050.16↑ higherSource ↗official
DystopiaBenchbasaglia_score#23 / 5060.9↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#27 / 5058.1↓ lowerSource ↗official
DystopiaBenchhuxley_score#27 / 5072.8↓ lowerSource ↗official
DystopiaBenchlaguardia_score#37 / 5069.03↓ lowerSource ↗official
DystopiaBenchorwell_score#28 / 5069.83↓ lowerSource ↗official
DystopiaBenchpetrov_score#30 / 5074.3↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#148 / 26014.73↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#246 / 26070.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#106 / 26076.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#204 / 25891.64↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#10 / 5428.57↑ higherSource ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#1 / 1099.2↑ higherSource ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#10 / 1032↑ higherSource ↗official
Inkling-Small model card — StrongREJECTsafety_rate#10 / 1097.4↑ higherSource ↗official
Manager Coercion Benchcoercion_ladder_depth#1 / 313.233↓ lowerSource ↗self run
MANTAAWMS#4 / 70.417↑ higherSource ↗official
MANTAAWVS#3 / 70.508↑ higherSource ↗official
PHAREbias_resistance_diagnostic#59 / 660.323↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#13 / 700.8188↑ higherSource ↗official
PHAREharm_resistance_diagnostic#28 / 700.9507↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#42 / 670.4549↑ higherSource ↗official
SM-Benchadversarial#36 / 7381.95↑ higherSource ↗official
SM-Benchambiguous_interpretation#60 / 7375.89↑ higherSource ↗official
SM-Benchanti_hallucination#30 / 7394.24↑ higherSource ↗official
SM-Bencheq_boundaries#17 / 7369.94↑ higherSource ↗official
SM-Benchoverfit#50 / 7359.56↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#30 / 1026.9↑ higherSource ↗official
SpeciEvalland_animal_4ns#67 / 1024.7↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#49 / 1024.75↓ lowerSource ↗official
SpeciEvalspeciesism#53 / 1022.05↓ lowerSource ↗official
TACbase_welfare_rate#53 / 6821.15↑ higherSource ↗official
ToolPrivacyBenchprivate_mt_poi#6 / 927.33↓ lowerSource ↗official
ToolPrivacyBenchpublic_mt_poi#6 / 918.99↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-16.6
Government46
Diplomacy66
Economy46.2
Society59

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression2.27
Traditional ↔ Secular-1.75