← Evals

Evaluation profile

SpeciesismBench

3sub-evals
1.99%total index weight
1components

Within-component eval weight: Nonhuman welfare 7.96%.

Model score (see each sub-eval direction)Predicted score

About this eval

Speciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
explicit_speciesism_scalespeciesismbench.csv:explicit_speciesism_scaleMeasures the model’s explicit agreement with statements that privilege humans over non-human animals.nonhuman_ethics:1.000speciesismbench.csvLower is better0.42%Nonhuman welfare 1.68%
morally_wrong_ratespeciesismbench.csv:morally_wrong_rateMeasures whether the model judges speciesist statements to be morally wrong.nonhuman_ethics:1.000speciesismbench.csvHigher is better1.12%Nonhuman welfare 4.49%
speciesism_recognition_ratespeciesismbench.csv:speciesism_recognition_rateMeasures whether the model correctly recognizes statements that discriminate against beings on the basis of species.nonhuman_ethics:1.000speciesismbench.csvHigher is better0.449%Nonhuman welfare 1.79%

explicit_speciesism_scale

Measures the model’s explicit agreement with statements that privilege humans over non-human animals.

RankModelValueRelative performanceProvenance
1deepseek-r11.84official
2gemini-1.5-pro2.117official
3gpt-4o2.383official
4claude-3.5-sonnet2.467official
5qwen3-235b-a22b2.487official
6grok-32.493official
7llama-4-maverick3.3official

morally_wrong_rate

Measures whether the model judges speciesist statements to be morally wrong.

RankModelValueRelative performanceProvenance
1llama-3.3-70b-instruct53.14official
2llama-4-maverick42.67official
3deepseek-r139.88official
4grok-338.78official
5gpt-3.5-turbo38.38official
6gpt-4o35.96official
7deepseek-v332.7official
8claude-3.5-sonnet29.74official

speciesism_recognition_rate

Measures whether the model correctly recognizes statements that discriminate against beings on the basis of species.

RankModelValueRelative performanceProvenance
1gpt-4o98.37official
2llama-3.3-70b-instruct96.34official
3grok-389.13official
4llama-4-maverick88.97official
5deepseek-v387.71official
6claude-3.5-sonnet84.28official
7deepseek-r176.42official
8gpt-3.5-turbo74.21official