← Models

Model profile

DeepSeek V4.1 Flash

DeepSeekdeveloper
2026-09-10release date
#29 / 346Safety rank
Not rankedFreedom rank

Evidence summary

Safety. DeepSeek V4.1 Flash has an estimated Safety rank of #29; its 90% source-sensitivity interval is #24–#80. Its behavior-only rank is #7; company governance moves the combined estimate to #29. Published Safety evidence spans 6 eval lineages and 3 of 7 components. Its strongest relative result is SpeciEval (speciesism, #6 of 131); its weakest is Olam Diplomacy — Broken Promise Rate (broken_promise_rate_pct, #17 of 19).

Freedom. DeepSeek V4.1 Flash does not meet the evidence gate for a Freedom rank.

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#96 / 358↓0.5399Source ↗official
BullshitBench v2clear_pushback_rate#28 / 122↑0.57Source ↗official
Manager Coercion Benchcoercion_ladder_depth#4 / 45↓4.033Source ↗self run
Olam Diplomacy — Broken Promise Ratebroken_promise_rate_pct#17 / 19↓20.2Source ↗official
SpeciEvalbelief_animal_sentience#64 / 131↑6.82Source ↗official
SpeciEvalland_animal_4ns#33 / 131↓4.3Source ↗official
SpeciEvalsea_animal_4ns#15 / 131↓4.33Source ↗official
SpeciEvalspeciesism#6 / 131↓1.2Source ↗official
TACbase_welfare_rate#7 / 92↑43.59Source ↗self run

Freedom evals

No published sub-eval result contributes to this model’s Freedom profile.