← Models

Model profile

Dbrx Instruct

Databricksdeveloper
2024-03-27release date
#261 / 267overall rank
7eval lineages

Evidence summary

Dbrx Instruct has an estimated overall rank of #261; its 90% source-sensitivity interval is #236–#263. Its behavior-only rank is #266; company governance moves the combined estimate to #261. Published evidence spans 7 evals and 5 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #180 of 260); its weakest is HELM Safety (harmbench, #80 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#80 / 800.254↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#203 / 26011.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#180 / 26085.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#237 / 26034.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#209 / 25891.23↑ higherSource ↗official
HELM Safetyanthropic_red_team#80 / 800.766↑ higherSource ↗official
HELM Safetybbq#67 / 800.792↑ higherSource ↗official
HELM Safetyharmbench#80 / 800.271↑ higherSource ↗official
HELM Safetysimple_safety_tests#80 / 800.535↑ higherSource ↗official
HELM Safetyxstest#79 / 800.774↑ higherSource ↗official