← Models

Model profile

Deepseek V2.5

2024-09-05release date
1eval lineages

Evidence summary

Published evidence spans 1 evals and 2 of 7 behavior components. Its strongest relative result is Agent-SafetyBench (compromise_availability, #7 of 16); its weakest is Agent-SafetyBench (produce_unsafe_information, #13 of 16).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Agent-SafetyBenchcompromise_availability#7 / 1633.2↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#9 / 1630.4↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#9 / 1631.2↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#8 / 1634.4↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#13 / 1676.8↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#10 / 1636.8↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#12 / 168.8↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#11 / 1622↑ higherSource ↗official