← Evals

Evaluation profile

TAC

1sub-evals
6.09%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Agentic animal-welfare behavior in travel-booking decisions under a neutral prompt.

Included in the behavior ranking.

Interpretation and limitations
  • MiMo-V2.6-Pro completed 155 of 156 purchases (99.36%) and has 55 of 156 welfare-preserving outcomes (35.26%). MiMo-V2.6-Flash completed 146 of 156 purchases (93.59%), below the existing 95% gate, and is excluded from active TAC scoring. Neither model's behavioral non-completions were resampled. Xiaomi native enabled reasoning is not asserted to be medium-effort equivalent.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
base_welfare_ratetac/tac-research-preview-round52.csv:base_welfare_rateMeasures whether an autonomous shopping agent preserves a stated animal-welfare constraint when making purchase decisions.Safety: nonhuman_ethics:1.000tacSafety: higher6.09%—

base_welfare_rate

Measures whether an autonomous shopping agent preserves a stated animal-welfare constraint when making purchase decisions.

RankModelValueRelative performanceProvenance
1claude-opus-4.864.74official
2claude-opus-559.62official
3claude-fable-555.77official
4qwen2.5-72b-instruct49.4self run
5qwen3-32b48.1self run
6qwen2.5-7b-instruct46.2self run
7deepseek-v4.1-flash43.59self run
7gpt-6-sol43.59self run
7muse-spark-1.343.59self run
10deepseek-v342.3self run
11granite-4.2-8b41.67self run
12gemini-2.5-flash39.74official
13claude-sonnet-4.637.82official
14gpt-5-nano37.82self run
15gpt-4.137.18official
16nemotron-3-nano-omni-30b-a3b36.54self run
17deepseek-v3.236.54official
18qwen3-235b-a22b35.9self run
19gemini-3-flash-preview35.9self run
20llama-3.3-70b-instruct35.3self run
21inkling-small35.26self run
21mimo-v2.6-pro35.26self run
21ministral-3-3b35.26self run
21nex-n2-mini35.26self run
25hy4-preview33.97self run
26glm-4.7-flash33.33self run
27deepseek-v3.132.1self run
28nemotron-3-nano-30b-a3b32.05self run
28solar-mini432.05self run
30seed-1.6-flash31.41self run
30seed-2.0-mini31.41self run
32gpt-oss-120b31.4self run
33gemini-3.8-flash30.77self run
34gemini-3.1-flash-lite30.13self run
35nemotron-3-super-120b-a12b29.5self run
36glm-5.329.49self run
36gpt-4.1-mini29.49self run
36ling-3.0-tiny29.49self run
39gpt-5.6-terra29.49official
40qwen3.5-9b28.85self run
40sakana-namazu28.85self run
42gemma-4-26b-a4b-it28.21self run
42gpt-oss-safeguard-20b28.21self run
42laguna-s-2.1-poolside28.21self run
42qwen3.5-397b-a17b28.21self run
46laguna-xs-2.1-poolside27.56self run
47gpt-5.226.28official
48big-pickle26.28self run
48gpt-5-mini26.28self run
48gpt-6-luna26.28self run
48solar-pro-426.28self run
52gemini-3.7-flash25.64official
53inkling25.64self run
53kimi-k2.625.64self run
55kimi-k325official
55space-bunny-alpha25self run
57glm-4.5-air24.36self run
57kimi-k2.524.36self run
57mercury-2.5-preview24.36self run
57minimax-m324.36self run
57qwen3.5-flash24.36self run
57qwen3.7-flash24.36self run
63gemini-3.5-flash-lite23.72self run
63gemma-4-31b-it23.72self run
63nemotron-3-ultra-550b-a55b23.72self run
66mimo-v2.5-pro23.7self run
67deepseek-v4-pro23.08official
68mimo-v2.521.79self run
69deepseek-v4-flash21.15official
70minimax-m2.721.15self run
70ox-alpha21.15self run
72deepseek-v4-flash-073120.51self run
72minimax-m2.520.51self run
74step-3.7-flash19.87self run
75glm-5.219.23official
75gpt-5.519.23official
77ling-2.6-flash19.23self run
77mercury-2.519.23self run
77seed-2.1-turbo19.23self run
80glm-518.59self run
80step-3.5-flash18.59self run
82glm-5.117.95self run
83gemini-3.5-flash17.31official
84hy316.67self run
84ornith-1.5-35b-a3b16.67self run
84qwen3.8-27b16.67self run
87gpt-5.6-sol16.67official
87grok-4.616.67official
87qwen3.8-max16.67official
90claude-sonnet-515.38official
91gpt-5.6-luna13.46self run
92muse-glimmer9.615self run