← Models

Model profile

Doubao 1.5 Thinking Pro

ByteDancedeveloper
2025-04-15release date
Not rankedSafety rank
#81 / 662Freedom rank

Evidence summary

Safety. Doubao 1.5 Thinking Pro does not meet the evidence gate for a Safety rank. Published Safety evidence spans 1 eval lineages and 4 of 7 components. Its strongest relative result is Concordia — Shutdown-Resistance (safety_score, #1 of 53); its weakest is Concordia — SciKnowEval-ChemicalHarmfulQA (safety_score, #63 of 63).

Freedom. Doubao 1.5 Thinking Pro has an estimated Freedom rank of #81; its 90% source-sensitivity interval is #15–#461. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — SciKnowEval-ChemicalHarmfulQA (safety_score, #1 of 63); its weakest is Concordia — SOSBench-Bio (safety_score, #7 of 63).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — Agentic-Misalignmentsafety_score#26 / 54↑80Source ↗official
Concordia — AirBench-Deceptionsafety_score#58 / 63↑46.67Source ↗official
Concordia — AirBench-Manipulationsafety_score#55 / 56↑50Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#61 / 63↑46.4Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#58 / 63↑68.13Source ↗official
Concordia — MASKsafety_score#48 / 62↑49.16Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#61 / 63↑7.407Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#63 / 63↑0.6608Source ↗official
Concordia — Shutdown-Resistancesafety_score#1 / 53↑100Source ↗official
Concordia — SOSBench-Biosafety_score#57 / 63↑19Source ↗official
Concordia — SOSBench-Chemsafety_score#58 / 63↑40.4Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.