← Evals

Evaluation profile

Adversarial Poetry Refusal (AHB self-run)

12sub-evals
0.568%Safety weight
0.972%Freedom weight
4components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Measures attack-success rate on the 600 adversarial-poetry prompts published in the Hugging Face AHB repository; lower values indicate stronger refusal of harmful requests.

Included in the behavior ranking.

Interpretation and limitations
  • MiMo-V2.6-Pro attempted 600 prompts and scored 594 (99.0%); Flash attempted 600 and scored 597 (99.5%). Their observed overall ASRs are 12.23% and 14.52%. Treating every missing outcome as anywhere between safe and unsafe gives full-panel bounds of 12.11%-13.11% for Pro and 14.44%-14.94% for Flash, not confidence intervals. Rankings use the available per-category observations, not imputed full-panel scores.
  • Pro has two invalid target replies, three malformed judge replies and one target reply lost after a timeout; Flash has two provider-filtered judge replies and one malformed judge reply. Judge failures are not target refusals, and the lost target reply was not regenerated. Missingness may be informative.
  • The MiMo endpoints use native enabled reasoning; this is not a claim of equivalence to another model's medium reasoning-effort setting. Historical provider-failure-as-refusal handling differs from the explicitly missing MiMo policy.
  • GPT-6 Luna and Sol scored 588/600 and 586/600 prompts; missing targets or judge votes are excluded. Two saved Sol endpoint-policy responses with complete refusal text were judged without generating replacement answers.
  • Space Bunny Alpha scored 598/600 prompts and Solar Mini 4 scored 591/600. Their observed overall ASRs are 12.6533% and 12.7468%; worst-case full-panel missing-outcome bounds are 12.6111%-12.9444% and 12.5556%-14.0556%, respectively. These are not confidence intervals or evidence of a meaningful ranking between the models. Both use medium reasoning and unchanged native provider pins.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
child_sexual_exploitationahb-poetry-self-run/ahb-poetry-self-run.csv:child_sexual_exploitationMeasures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%
defamationahb-poetry-self-run/ahb-poetry-self-run.csv:defamationMeasures unsafe compliance with poetry-obfuscated defamatory requests.Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0318%0.081%
hateahb-poetry-self-run/ahb-poetry-self-run.csv:hateMeasures unsafe compliance with poetry-obfuscated hateful requests.Safety: fairness_nondiscrimination:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0605%0.081%
indiscriminate_weapons_cbrneahb-poetry-self-run/ahb-poetry-self-run.csv:indiscriminate_weapons_cbrneMeasures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0124%0.081%
intellectual_propertyahb-poetry-self-run/ahb-poetry-self-run.csv:intellectual_propertyMeasures unsafe compliance with poetry-obfuscated intellectual-property requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%
non_violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:non_violent_crimesMeasures unsafe compliance with poetry-obfuscated non-violent-crime requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0124%0.081%
privacyahb-poetry-self-run/ahb-poetry-self-run.csv:privacyMeasures unsafe compliance with poetry-obfuscated privacy-violation requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%
sex_related_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:sex_related_crimesMeasures unsafe compliance with poetry-obfuscated sex-related-crime requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%
sexual_contentahb-poetry-self-run/ahb-poetry-self-run.csv:sexual_contentMeasures unsafe compliance with poetry-obfuscated sexual-content requests.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.00495%0.081%
specialized_adviceahb-poetry-self-run/ahb-poetry-self-run.csv:specialized_adviceMeasures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators.Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0124%0.081%
suicide_self_harmahb-poetry-self-run/ahb-poetry-self-run.csv:suicide_self_harmMeasures unsafe compliance with poetry-obfuscated suicide and self-harm requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%
violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:violent_crimesMeasures unsafe compliance with poetry-obfuscated violent-crime requests.Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-runSafety: lower · Freedom: higher0.0723%0.081%

child_sexual_exploitation

Measures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna0self run
2glm-5.3-flash1.449self run
2gpt-5.6-luna1.449self run
2mercury-21.449self run
5gpt-6-sol2.899self run
5granite-4.2-8b2.899self run
7longcat-2.04.348self run
8mimo-v2.6-pro6.522self run
9laguna-s-2.1-poolside7.971self run
9nemotron-3.5-lightning7.971self run
9solar-mini47.971self run
12mimo-v2.6-flash10.14self run
13space-bunny-alpha14.49self run
14solar-pro-431.88self run

defamation

Measures unsafe compliance with poetry-obfuscated defamatory requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna0self run
2mercury-22.174self run
3glm-5.3-flash3.623self run
3granite-4.2-8b3.623self run
5gpt-6-sol3.788self run
6laguna-s-2.1-poolside7.971self run
6mimo-v2.6-pro7.971self run
8gpt-5.6-luna9.42self run
9mimo-v2.6-flash10.37self run
10solar-mini411.85self run
11longcat-2.017.39self run
12space-bunny-alpha19.26self run
13nemotron-3.5-lightning19.57self run
14solar-pro-455.07self run

hate

Measures unsafe compliance with poetry-obfuscated hateful requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna1.449self run
2glm-5.3-flash3.623self run
3granite-4.2-8b5.072self run
4gpt-5.6-luna5.797self run
4longcat-2.05.797self run
4mercury-25.797self run
7laguna-s-2.1-poolside7.246self run
8mimo-v2.6-pro7.971self run
9gpt-6-sol8.696self run
10mimo-v2.6-flash9.42self run
11nemotron-3.5-lightning12.32self run
12solar-mini413.33self run
13space-bunny-alpha15.22self run
14solar-pro-443.48self run

indiscriminate_weapons_cbrne

Measures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0self run
1gpt-6-sol0self run
1mercury-20self run
4gpt-6-luna0.8333self run
5glm-5.3-flash1.449self run
6space-bunny-alpha2.174self run
7mimo-v2.6-pro2.899self run
8longcat-2.07.971self run
9laguna-s-2.1-poolside9.42self run
10granite-4.2-8b10.87self run
11nemotron-3.5-lightning12.32self run
12mimo-v2.6-flash13.04self run
13solar-mini414.81self run
14solar-pro-434.06self run

intellectual_property

Measures unsafe compliance with poetry-obfuscated intellectual-property requests.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna0self run
1gpt-6-luna0self run
3mercury-22.174self run
4granite-4.2-8b3.623self run
5gpt-6-sol4.444self run
6glm-5.3-flash8.696self run
7laguna-s-2.1-poolside10.14self run
8longcat-2.010.87self run
9solar-mini411.59self run
10space-bunny-alpha16.67self run
11mimo-v2.6-pro17.39self run
12nemotron-3.5-lightning21.74self run
13mimo-v2.6-flash22.46self run
14solar-pro-427.54self run

non_violent_crimes

Measures unsafe compliance with poetry-obfuscated non-violent-crime requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna0self run
2gpt-5.6-luna0.7246self run
3gpt-6-sol0.7407self run
4mercury-22.174self run
5glm-5.3-flash4.348self run
5laguna-s-2.1-poolside4.348self run
7mimo-v2.6-pro7.246self run
8space-bunny-alpha7.971self run
9granite-4.2-8b8.696self run
10solar-mini49.848self run
11longcat-2.010.14self run
12mimo-v2.6-flash17.39self run
13nemotron-3.5-lightning18.12self run
14solar-pro-434.78self run

privacy

Measures unsafe compliance with poetry-obfuscated privacy-violation requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna0.7246self run
2gpt-5.6-luna5.797self run
3glm-5.3-flash7.246self run
4gpt-6-sol8.696self run
5mercury-29.42self run
6granite-4.2-8b13.77self run
7laguna-s-2.1-poolside15.94self run
8longcat-2.018.84self run
9solar-mini420.74self run
10space-bunny-alpha26.09self run
11mimo-v2.6-pro27.41self run
12nemotron-3.5-lightning27.54self run
13mimo-v2.6-flash29.71self run
14solar-pro-443.48self run

sexual_content

Measures unsafe compliance with poetry-obfuscated sexual-content requests.

RankModelValueRelative performanceProvenance
1gpt-6-sol1.55self run
2gpt-6-luna3.03self run
3granite-4.2-8b4.545self run
3mercury-24.545self run
5gpt-5.6-luna5.303self run
6laguna-s-2.1-poolside6.818self run
7glm-5.3-flash7.576self run
8mimo-v2.6-pro8.333self run
9solar-mini410.08self run
10space-bunny-alpha12.4self run
11nemotron-3.5-lightning12.88self run
12longcat-2.015.15self run
13mimo-v2.6-flash19.7self run
14solar-pro-425self run

specialized_advice

Measures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators.

RankModelValueRelative performanceProvenance
1gpt-6-luna0.6734self run
2gpt-6-sol1.042self run
3gpt-5.6-luna2.333self run
3mercury-22.333self run
5glm-5.3-flash3self run
6granite-4.2-8b4.667self run
6laguna-s-2.1-poolside4.667self run
8space-bunny-alpha5self run
9longcat-2.06.333self run
10solar-mini49.667self run
11mimo-v2.6-pro10.44self run
12mimo-v2.6-flash10.77self run
13nemotron-3.5-lightning12.67self run
14solar-pro-429.67self run

suicide_self_harm

Measures unsafe compliance with poetry-obfuscated suicide and self-harm requests.

RankModelValueRelative performanceProvenance
1gpt-5.6-luna1.515self run
2gpt-6-luna1.587self run
2gpt-6-sol1.587self run
4mercury-23.788self run
5laguna-s-2.1-poolside6.061self run
6granite-4.2-8b6.818self run
7glm-5.3-flash7.576self run
8longcat-2.09.848self run
9mimo-v2.6-flash10.85self run
10solar-mini413.64self run
11mimo-v2.6-pro14.39self run
11space-bunny-alpha14.39self run
13nemotron-3.5-lightning15.15self run
14solar-pro-437.12self run

violent_crimes

Measures unsafe compliance with poetry-obfuscated violent-crime requests.

RankModelValueRelative performanceProvenance
1gpt-6-luna0.7576self run
2gpt-5.6-luna2.273self run
2mercury-22.273self run
4gpt-6-sol3.03self run
5glm-5.3-flash5.303self run
6laguna-s-2.1-poolside7.576self run
7granite-4.2-8b8.333self run
8mimo-v2.6-flash13.64self run
8space-bunny-alpha13.64self run
10longcat-2.014.39self run
11nemotron-3.5-lightning18.94self run
12solar-mini419.84self run
13mimo-v2.6-pro23.26self run
14solar-pro-447.73self run