← Models

Model profile

DeepSeek v3.1 Terminus

DeepSeekdeveloper
2025-09-22release date
#245 / 346Safety rank
#50 / 662Freedom rank

Evidence summary

Safety. DeepSeek v3.1 Terminus has an estimated Safety rank of #245; its 90% source-sensitivity interval is #110–#306. Its behavior-only rank is #236; company governance moves the combined estimate to #245. Published Safety evidence spans 4 eval lineages and 4 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is Concordia — FRT-AirBench-SecurityRisks (safety_score, #44 of 45).

Freedom. DeepSeek v3.1 Terminus has an estimated Freedom rank of #50; its 90% source-sensitivity interval is #69–#296. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — FRT-SciKnowEval-BiologicalHarmfulQA (safety_score, #2 of 45); its weakest is Concordia — SciKnowEval-BiologicalHarmfulQA (safety_score, #31 of 63).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#153 / 358↓0.7483Source ↗official
AgentDrive Safety Compliancescr#20 / 48↑88.75Source ↗official
Concordia — Agentic-Misalignmentsafety_score#48 / 54↑52.17Source ↗official
Concordia — AirBench-Deceptionsafety_score#40 / 63↑75.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#38 / 56↑79.33Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#45 / 63↑82.66Source ↗official
Concordia — APEsafety_score#47 / 55↑3.255Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#47 / 63↑76.1Source ↗official
Concordia — DarkBenchsafety_score#25 / 55↑57.51Source ↗official
Concordia — Fortress-Biologicalsafety_score#52 / 54↑23Source ↗official
Concordia — Fortress-Chemicalsafety_score#50 / 54↑28.42Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#39 / 54↑45.29Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#41 / 45↑15.11Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#44 / 45↑7.833Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#42 / 45↑0.3333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#34 / 45↑35.67Source ↗official
Concordia — MASKsafety_score#58 / 62↑40.94Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#32 / 63↑82.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#41 / 63↑14.57Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53↑100Source ↗official
Concordia — SOSBench-Biosafety_score#51 / 63↑50.9Source ↗official
Concordia — SOSBench-Chemsafety_score#48 / 63↑69.4Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#9 / 27↑0.72Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#24 / 63↓75.19Source ↗official
Concordia — AirBench-Manipulationsafety_score#18 / 56↓79.33Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#19 / 56↓49.52Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#18 / 63↓82.66Source ↗official
Concordia — Fortress-Biologicalsafety_score#3 / 54↓23Source ↗official
Concordia — Fortress-Chemicalsafety_score#5 / 54↓28.42Source ↗official
Concordia — Fortress-Privacy/Scamssafety_score#16 / 54↓45.29Source ↗official
Concordia — FRT-AirBench-Manipulationsafety_score#4 / 45↓15.11Source ↗official
Concordia — FRT-AirBench-SecurityRiskssafety_score#2 / 45↓7.833Source ↗official
Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score#2 / 45↓0.3333Source ↗official
Concordia — FRT-SOSBench-Chemsafety_score#12 / 45↓35.67Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#31 / 63↓82.49Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#23 / 63↓14.57Source ↗official
Concordia — SOSBench-Biosafety_score#13 / 63↓50.9Source ↗official
Concordia — SOSBench-Chemsafety_score#16 / 63↓69.4Source ↗official
SpeechMap model completioncomplete_pct#61 / 181↑68.6Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#33 / 156↑5.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#46 / 156↑4Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-23.4
Government48
Diplomacy67.8
Economy44.6
Society64.9