← Evals

Evaluation profile

DecodingTrust

3sub-evals
0.24%total index weight
2components

Within-component eval weight: Misuse resistance 0.528% · Fairness 1.87%.

Model score (higher is better)Predicted score

About this eval

Multidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
machine_ethicsdecodingtrust/scoring-dimensions.csv:machine_ethicsMeasures whether the model selects ethically appropriate actions in machine-ethics scenarios.ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csvHigher is better0.0151%Misuse resistance 0.151%
stereotype_biasdecodingtrust/scoring-dimensions.csv:stereotype_biasMeasures whether the model avoids reproducing demographic stereotypes across bias prompts.fairness_nondiscrimination:1.000decodingtrust/scoring-dimensions.csvHigher is better0.187%Fairness 1.87%
toxicitydecodingtrust/scoring-dimensions.csv:toxicityMeasures whether the model avoids generating toxic continuations under challenging prompt conditions.ordinary_harm_misuse_resistance:1.000decodingtrust/scoring-dimensions.csvHigher is better0.0377%Misuse resistance 0.377%

machine_ethics

Measures whether the model selects ethically appropriate actions in machine-ethics scenarios.

RankModelValueRelative performanceProvenance
1gemini-1.0-pro93.74official
2gpt-3.5-turbo86.38official
3claude-285.17official
4gpt-476.6official
5tulu-2-7b49official
6zephyr-7b-beta41.03official
7llama-2-7b-chat40.58official
8tulu-2-13b36.64official

stereotype_bias

Measures whether the model avoids reproducing demographic stereotypes across bias prompts.

RankModelValueRelative performanceProvenance
1claude-2100official
2gemini-1.0-pro98.33official
3llama-2-7b-chat97.6official
4tulu-2-7b96.6official
5zephyr-7b-beta92.6official
6tulu-2-13b89.33official
7gpt-3.5-turbo87official
8gpt-477official

toxicity

Measures whether the model avoids generating toxic continuations under challenging prompt conditions.

RankModelValueRelative performanceProvenance
1claude-292.11official
2llama-2-7b-chat80official
3gemini-1.0-pro77.53official
4gpt-3.5-turbo47official
5tulu-2-13b44.8official
6gpt-441official
7zephyr-7b-beta31.97official
8tulu-2-7b29.46official