Evaluation profile
TAC
1sub-evals
5.89%total index weight
1components
Within-component eval weight: Nonhuman welfare 23.5%.
Model score (higher is better)Predicted score
About this eval
Agentic animal-welfare behavior in travel-booking decisions under a neutral prompt.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| base_welfare_ratetac/tac-research-preview-round52.csv:base_welfare_rateMeasures whether an autonomous shopping agent preserves a stated animal-welfare constraint when making purchase decisions. | nonhuman_ethics:1.000tac | Higher is better | 5.89% | Nonhuman welfare 23.5% |
base_welfare_rate
Measures whether an autonomous shopping agent preserves a stated animal-welfare constraint when making purchase decisions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.8 | 64.74 | official | |
| 2 | claude-opus-5 | 59.62 | official | |
| 3 | claude-fable-5 | 55.77 | official | |
| 4 | qwen2.5-72b-instruct | 49.4 | self run | |
| 5 | qwen3-32b | 48.1 | self run | |
| 6 | qwen-2.5-7b-instruct | 46.2 | self run | |
| 7 | deepseek-v3 | 42.3 | self run | |
| 8 | gemini-2.5-flash | 39.74 | official | |
| 9 | claude-sonnet-4.6 | 37.82 | official | |
| 10 | gpt-5-nano | 37.82 | self run | |
| 11 | gpt-4.1 | 37.18 | official | |
| 12 | nemotron-3-nano-omni-30b-a3b | 36.54 | self run | |
| 13 | deepseek-v3.2 | 36.54 | official | |
| 14 | qwen3-235b-a22b | 35.9 | self run | |
| 15 | gemini-3-flash-preview | 35.9 | self run | |
| 16 | llama-3.3-70b-instruct | 35.3 | self run | |
| 17 | inkling-small | 35.26 | self run | |
| 17 | ministral-3b-2512 | 35.26 | self run | |
| 17 | nex-n2-mini | 35.26 | self run | |
| 20 | glm-4.7-flash | 33.33 | self run | |
| 21 | deepseek-v3.1 | 32.1 | self run | |
| 22 | nemotron-3-nano-30b-a3b | 32.05 | self run | |
| 23 | seed-1.6-flash | 31.41 | self run | |
| 23 | seed-2.0-mini | 31.41 | self run | |
| 25 | gpt-oss-120b | 31.4 | self run | |
| 26 | gemini-3.1-flash-lite | 30.13 | self run | |
| 27 | nemotron-3-super-120b-a12b | 29.5 | self run | |
| 28 | gpt-4.1-mini | 29.49 | self run | |
| 29 | gpt-5.6-terra | 29.49 | official | |
| 30 | qwen3.5-9b | 28.85 | self run | |
| 31 | gemma-4-26b-a4b-it | 28.21 | self run | |
| 31 | gpt-oss-safeguard-20b | 28.21 | self run | |
| 31 | laguna-s-2.1-poolside | 28.21 | self run | |
| 31 | qwen3.5-397b-a17b | 28.21 | self run | |
| 35 | laguna-xs-2.1-poolside | 27.56 | self run | |
| 36 | gpt-5.2 | 26.28 | official | |
| 37 | big-pickle | 26.28 | self run | |
| 37 | gpt-5-mini | 26.28 | self run | |
| 39 | inkling | 25.64 | self run | |
| 39 | kimi-k2.6 | 25.64 | self run | |
| 41 | kimi-k3 | 25 | official | |
| 42 | glm-4.5-air | 24.36 | self run | |
| 42 | kimi-k2.5 | 24.36 | self run | |
| 42 | minimax-m3 | 24.36 | self run | |
| 42 | qwen3.5-flash-02-23 | 24.36 | self run | |
| 42 | qwen3.7-flash | 24.36 | self run | |
| 47 | gemini-3.5-flash-lite | 23.72 | self run | |
| 47 | gemma-4-31b-it | 23.72 | self run | |
| 47 | nemotron-3-ultra-550b-a55b | 23.72 | self run | |
| 50 | mimo-v2.5-pro | 23.7 | self run | |
| 51 | deepseek-v4-pro | 23.08 | official | |
| 52 | mimo-v2.5 | 21.79 | self run | |
| 53 | deepseek-v4-flash | 21.15 | official | |
| 54 | minimax-m2.7 | 21.15 | self run | |
| 55 | minimax-m2.5 | 20.51 | self run | |
| 56 | step-3.7-flash | 19.87 | self run | |
| 57 | glm-5.2 | 19.23 | official | |
| 57 | gpt-5.5 | 19.23 | official | |
| 59 | ling-2.6-flash | 19.23 | self run | |
| 60 | glm-5 | 18.59 | self run | |
| 60 | step-3.5-flash | 18.59 | self run | |
| 62 | glm-5.1 | 17.95 | self run | |
| 63 | gemini-3.5-flash | 17.31 | official | |
| 64 | hy3 | 16.67 | self run | |
| 65 | gpt-5.6-sol | 16.67 | official | |
| 65 | qwen3.8-max | 16.67 | official | |
| 67 | claude-sonnet-5 | 15.38 | official | |
| 68 | gpt-5.6-luna | 13.46 | self run |