← Models

Model profile

Mixtral 8x7B Instruct

Mistral AIdeveloper
2023-12-10release date
#327 / 346Safety rank
#58 / 662Freedom rank

Evidence summary

Safety. Mixtral 8x7B Instruct has an estimated Safety rank of #327; its 90% source-sensitivity interval is #245–#341. Its behavior-only rank is #317; company governance moves the combined estimate to #327. Published Safety evidence spans 13 eval lineages and 6 of 7 components. Its strongest relative result is SALAD-Bench (mcq_information_safety_harms, #6 of 33); its weakest is HELM Safety (harmbench, #77 of 80).

Freedom. Mixtral 8x7B Instruct has an estimated Freedom rank of #58; its 90% source-sensitivity interval is #64–#215. Published Freedom evidence spans 12 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is HELM Safety (xstest, #61 of 80).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#74 / 80↑0.391Source ↗official
CASE-Benchagreement_accuracy#4 / 7↑83.11Source ↗official
COMPL-AI AI-Identity Disclosurescore#9 / 14↑0.8904Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#13 / 14↑0.2561Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#12 / 13↑0.375Source ↗official
Contextual MoralChoicehuman_agreement#6 / 22↑0.45Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#221 / 270↑11.37Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#147 / 270↑86.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#240 / 270↑38.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#175 / 268↑95.18Source ↗official
HarmBenchdr#25 / 28↓47.3Source ↗official
HELM Safetyanthropic_red_team#75 / 80↑0.928Source ↗official
HELM Safetybbq#63 / 80↑0.857Source ↗official
HELM Safetyharmbench#77 / 80↑0.451Source ↗official
HELM Safetysimple_safety_tests#75 / 80↑0.905Source ↗official
HELM Safetyxstest#61 / 80↑0.931Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
MedSafetyBenchmedical_safety_score#9 / 30↑94Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#19 / 33↑11.85Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#25 / 33↑5.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#25 / 33↑7.67Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#18 / 33↑9.54Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#19 / 33↑10.35Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#21 / 33↑9.52Source ↗official
SALAD-Benchbase_human_autonomy_integrity#28 / 33↑76Source ↗official
SALAD-Benchbase_information_safety_harms#27 / 33↑86.8Source ↗official
SALAD-Benchbase_malicious_use#28 / 33↑67.65Source ↗official
SALAD-Benchbase_misinformation_harms#28 / 33↑84.39Source ↗official
SALAD-Benchbase_representation_toxicity#25 / 33↑82.05Source ↗official
SALAD-Benchbase_socioeconomic_harms#27 / 33↑80.85Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#9 / 33↑51.39Source ↗official
SALAD-Benchmcq_information_safety_harms#6 / 33↑51.94Source ↗official
SALAD-Benchmcq_malicious_use#9 / 33↑53.27Source ↗official
SALAD-Benchmcq_misinformation_harms#9 / 33↑52.86Source ↗official
SALAD-Benchmcq_representation_toxicity#9 / 33↑52.08Source ↗official
SALAD-Benchmcq_socioeconomic_harms#8 / 33↑48.89Source ↗official
SORRY-Benchavg#43 / 51↓0.56Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#7 / 80↓0.391Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#2 / 14↓0.2561Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#122 / 270↓86.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#27 / 270↓38.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#94 / 268↓95.18Source ↗official
HarmBenchdr#4 / 28↑47.3Source ↗official
HELM Safetyanthropic_red_team#6 / 80↓0.928Source ↗official
HELM Safetyharmbench#4 / 80↓0.451Source ↗official
HELM Safetysimple_safety_tests#6 / 80↓0.905Source ↗official
HELM Safetyxstest#61 / 80↑0.931Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
MedSafetyBenchmedical_safety_score#22 / 30↓94Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#15 / 33↓11.85Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#7 / 33↓5.21Source ↗official
SALAD-Benchattack_enhanced_malicious_use#9 / 33↓7.67Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#16 / 33↓9.54Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#15 / 33↓10.35Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#13 / 33↓9.52Source ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 33↓76Source ↗official
SALAD-Benchbase_information_safety_harms#7 / 33↓86.8Source ↗official
SALAD-Benchbase_malicious_use#6 / 33↓67.65Source ↗official
SALAD-Benchbase_misinformation_harms#6 / 33↓84.39Source ↗official
SALAD-Benchbase_representation_toxicity#9 / 33↓82.05Source ↗official
SALAD-Benchbase_socioeconomic_harms#7 / 33↓80.85Source ↗official
SORRY-Benchavg#8 / 51↑0.56Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#33 / 156↑5.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-5.8
Government45.9
Diplomacy57.2
Economy40.6
Society57.8