← Evals

Evaluation profile

HELM Classic RealToxicityPrompts

2sub-evals
0.207%total index weight
1components

Within-component eval weight: Misuse resistance 2.07%.

Model score (lower is better)Predicted score

About this eval

Toxic continuation generation under toxic and non-toxic prompt strata.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
non_toxic_prompt_toxic_fractionrealtoxicityprompts/realtoxicityprompts.csv:non_toxic_prompt_toxic_fractionMeasures whether the model generates toxic continuations when the prompt does not contain toxic language.ordinary_harm_misuse_resistance:1.000realtoxicitypromptsLower is better0.104%Misuse resistance 1.04%
toxic_prompt_toxic_fractionrealtoxicityprompts/realtoxicityprompts.csv:toxic_prompt_toxic_fractionMeasures whether the model generates toxic continuations when the prompt itself contains toxic language.ordinary_harm_misuse_resistance:1.000realtoxicitypromptsLower is better0.104%Misuse resistance 1.04%

non_toxic_prompt_toxic_fraction

Measures whether the model generates toxic continuations when the prompt does not contain toxic language.

RankModelValueRelative performanceProvenance
1xlarge-202206090.001287official
2command-medium-beta0.002317official
2medium-202207200.002317official
4large-202207200.003089official
4small-202207200.003089official
6palmyra-x0.003089official
7xlarge-202211080.003346official
8command-xlarge-beta0.003604official
8medium-202211080.003604official
10t5-11b0.004118official
11t0pp0.009266official
12text-davinci-0030.01004official
13luminous-base0.01055official
14luminous-extended0.01313official
15ul20.01364official
16text-ada-0010.01519official
17j1-large0.01725official
18j1-grande0.0175official
19gpt-j-6b0.01905official
20bloom0.01905official
21text-babbage-0010.01931official
22ada0.01982official
23babbage0.02059official
24j2-large0.02059official
25j1-jumbo0.02111official
26tnlgv2-7b0.02136official
27glm0.02188official
28gpt-neox-20b0.02188official
29text-curie-0010.02214official
30curie0.02291official
31j1-grande-v2-beta0.02342official
32opt-66b0.02368official
33opt-175b0.0242official
34j2-jumbo0.0242official
35stanford-online-all-v4-s30.02445official
36palmyra-instruct-300.02471official
36text-davinci-0020.02471official
38j2-grande0.02523official
39tnlgv2-530b0.02523official
40luminous-supreme0.02728official
41yalm0.02754official
42davinci0.03449official

toxic_prompt_toxic_fraction

Measures whether the model generates toxic continuations when the prompt itself contains toxic language.

RankModelValueRelative performanceProvenance
1palmyra-x0.008072official
2t5-11b0.01435official
3medium-202207200.01435official
4command-xlarge-beta0.01614official
4medium-202211080.01614official
6small-202207200.01704official
7command-medium-beta0.01794official
8large-202207200.01794official
9xlarge-202206090.01883official
10xlarge-202211080.02152official
11luminous-base0.04126official
12luminous-extended0.04305official
13text-ada-0010.06457official
14text-davinci-0030.06726official
15palmyra-instruct-300.07265official
15ul20.07265official
17gpt-neox-20b0.07803official
18opt-175b0.08072official
19gpt-j-6b0.0852official
20luminous-supreme0.0852official
21babbage0.0861official
22j1-large0.0861official
23ada0.08789official
24curie0.08969official
25tnlgv2-530b0.09148official
26j1-jumbo0.09148official
27opt-66b0.09238official
28stanford-online-all-v4-s30.09327official
29j1-grande0.09417official
30bloom0.09507official
31glm0.09507official
32tnlgv2-7b0.09596official
33j2-large0.09865official
34j2-jumbo0.09955official
35text-davinci-0020.1013official
36text-babbage-0010.1045official
37davinci0.1049official
37j1-grande-v2-beta0.1049official
39text-curie-0010.1067official
40j2-grande0.1094official
41t0pp0.1184official
42yalm0.157official