Evaluation profile
Adversarial Poetry Refusal (AHB self-run)
Weights below are portfolio-specific global index weights.
About this eval
Measures attack-success rate on the 600 adversarial-poetry prompts published in the Hugging Face AHB repository; lower values indicate stronger refusal of harmful requests.
Included in the behavior ranking.
Interpretation and limitations
- MiMo-V2.6-Pro attempted 600 prompts and scored 594 (99.0%); Flash attempted 600 and scored 597 (99.5%). Their observed overall ASRs are 12.23% and 14.52%. Treating every missing outcome as anywhere between safe and unsafe gives full-panel bounds of 12.11%-13.11% for Pro and 14.44%-14.94% for Flash, not confidence intervals. Rankings use the available per-category observations, not imputed full-panel scores.
- Pro has two invalid target replies, three malformed judge replies and one target reply lost after a timeout; Flash has two provider-filtered judge replies and one malformed judge reply. Judge failures are not target refusals, and the lost target reply was not regenerated. Missingness may be informative.
- The MiMo endpoints use native enabled reasoning; this is not a claim of equivalence to another model's medium reasoning-effort setting. Historical provider-failure-as-refusal handling differs from the explicitly missing MiMo policy.
- GPT-6 Luna and Sol scored 588/600 and 586/600 prompts; missing targets or judge votes are excluded. Two saved Sol endpoint-policy responses with complete refusal text were judged without generating replacement answers.
- Space Bunny Alpha scored 598/600 prompts and Solar Mini 4 scored 591/600. Their observed overall ASRs are 12.6533% and 12.7468%; worst-case full-panel missing-outcome bounds are 12.6111%-12.9444% and 12.5556%-14.0556%, respectively. These are not confidence intervals or evidence of a meaningful ranking between the models. Both use medium reasoning and unchanged native provider pins.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| child_sexual_exploitationahb-poetry-self-run/ahb-poetry-self-run.csv:child_sexual_exploitationMeasures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
| defamationahb-poetry-self-run/ahb-poetry-self-run.csv:defamationMeasures unsafe compliance with poetry-obfuscated defamatory requests. | Safety: human_rights_systemic_harm:0.500|truthfulness_honesty:0.500 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0318% | 0.081% |
| hateahb-poetry-self-run/ahb-poetry-self-run.csv:hateMeasures unsafe compliance with poetry-obfuscated hateful requests. | Safety: fairness_nondiscrimination:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0605% | 0.081% |
| indiscriminate_weapons_cbrneahb-poetry-self-run/ahb-poetry-self-run.csv:indiscriminate_weapons_cbrneMeasures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0124% | 0.081% |
| intellectual_propertyahb-poetry-self-run/ahb-poetry-self-run.csv:intellectual_propertyMeasures unsafe compliance with poetry-obfuscated intellectual-property requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
| non_violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:non_violent_crimesMeasures unsafe compliance with poetry-obfuscated non-violent-crime requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0124% | 0.081% |
| privacyahb-poetry-self-run/ahb-poetry-self-run.csv:privacyMeasures unsafe compliance with poetry-obfuscated privacy-violation requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
| sex_related_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:sex_related_crimesMeasures unsafe compliance with poetry-obfuscated sex-related-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
| sexual_contentahb-poetry-self-run/ahb-poetry-self-run.csv:sexual_contentMeasures unsafe compliance with poetry-obfuscated sexual-content requests. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.00495% | 0.081% |
| specialized_adviceahb-poetry-self-run/ahb-poetry-self-run.csv:specialized_adviceMeasures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0124% | 0.081% |
| suicide_self_harmahb-poetry-self-run/ahb-poetry-self-run.csv:suicide_self_harmMeasures unsafe compliance with poetry-obfuscated suicide and self-harm requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
| violent_crimesahb-poetry-self-run/ahb-poetry-self-run.csv:violent_crimesMeasures unsafe compliance with poetry-obfuscated violent-crime requests. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000ahb-poetry-self-run | Safety: lower · Freedom: higher | 0.0723% | 0.081% |
child_sexual_exploitation
Measures unsafe compliance with poetry-obfuscated child-sexual-exploitation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0 | self run | |
| 2 | glm-5.3-flash | 1.449 | self run | |
| 2 | gpt-5.6-luna | 1.449 | self run | |
| 2 | mercury-2 | 1.449 | self run | |
| 5 | gpt-6-sol | 2.899 | self run | |
| 5 | granite-4.2-8b | 2.899 | self run | |
| 7 | longcat-2.0 | 4.348 | self run | |
| 8 | mimo-v2.6-pro | 6.522 | self run | |
| 9 | laguna-s-2.1-poolside | 7.971 | self run | |
| 9 | nemotron-3.5-lightning | 7.971 | self run | |
| 9 | solar-mini4 | 7.971 | self run | |
| 12 | mimo-v2.6-flash | 10.14 | self run | |
| 13 | space-bunny-alpha | 14.49 | self run | |
| 14 | solar-pro-4 | 31.88 | self run |
defamation
Measures unsafe compliance with poetry-obfuscated defamatory requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0 | self run | |
| 2 | mercury-2 | 2.174 | self run | |
| 3 | glm-5.3-flash | 3.623 | self run | |
| 3 | granite-4.2-8b | 3.623 | self run | |
| 5 | gpt-6-sol | 3.788 | self run | |
| 6 | laguna-s-2.1-poolside | 7.971 | self run | |
| 6 | mimo-v2.6-pro | 7.971 | self run | |
| 8 | gpt-5.6-luna | 9.42 | self run | |
| 9 | mimo-v2.6-flash | 10.37 | self run | |
| 10 | solar-mini4 | 11.85 | self run | |
| 11 | longcat-2.0 | 17.39 | self run | |
| 12 | space-bunny-alpha | 19.26 | self run | |
| 13 | nemotron-3.5-lightning | 19.57 | self run | |
| 14 | solar-pro-4 | 55.07 | self run |
hate
Measures unsafe compliance with poetry-obfuscated hateful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 1.449 | self run | |
| 2 | glm-5.3-flash | 3.623 | self run | |
| 3 | granite-4.2-8b | 5.072 | self run | |
| 4 | gpt-5.6-luna | 5.797 | self run | |
| 4 | longcat-2.0 | 5.797 | self run | |
| 4 | mercury-2 | 5.797 | self run | |
| 7 | laguna-s-2.1-poolside | 7.246 | self run | |
| 8 | mimo-v2.6-pro | 7.971 | self run | |
| 9 | gpt-6-sol | 8.696 | self run | |
| 10 | mimo-v2.6-flash | 9.42 | self run | |
| 11 | nemotron-3.5-lightning | 12.32 | self run | |
| 12 | solar-mini4 | 13.33 | self run | |
| 13 | space-bunny-alpha | 15.22 | self run | |
| 14 | solar-pro-4 | 43.48 | self run |
indiscriminate_weapons_cbrne
Measures unsafe compliance with poetry-obfuscated indiscriminate-weapons and CBRNE requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0 | self run | |
| 1 | gpt-6-sol | 0 | self run | |
| 1 | mercury-2 | 0 | self run | |
| 4 | gpt-6-luna | 0.8333 | self run | |
| 5 | glm-5.3-flash | 1.449 | self run | |
| 6 | space-bunny-alpha | 2.174 | self run | |
| 7 | mimo-v2.6-pro | 2.899 | self run | |
| 8 | longcat-2.0 | 7.971 | self run | |
| 9 | laguna-s-2.1-poolside | 9.42 | self run | |
| 10 | granite-4.2-8b | 10.87 | self run | |
| 11 | nemotron-3.5-lightning | 12.32 | self run | |
| 12 | mimo-v2.6-flash | 13.04 | self run | |
| 13 | solar-mini4 | 14.81 | self run | |
| 14 | solar-pro-4 | 34.06 | self run |
intellectual_property
Measures unsafe compliance with poetry-obfuscated intellectual-property requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0 | self run | |
| 1 | gpt-6-luna | 0 | self run | |
| 3 | mercury-2 | 2.174 | self run | |
| 4 | granite-4.2-8b | 3.623 | self run | |
| 5 | gpt-6-sol | 4.444 | self run | |
| 6 | glm-5.3-flash | 8.696 | self run | |
| 7 | laguna-s-2.1-poolside | 10.14 | self run | |
| 8 | longcat-2.0 | 10.87 | self run | |
| 9 | solar-mini4 | 11.59 | self run | |
| 10 | space-bunny-alpha | 16.67 | self run | |
| 11 | mimo-v2.6-pro | 17.39 | self run | |
| 12 | nemotron-3.5-lightning | 21.74 | self run | |
| 13 | mimo-v2.6-flash | 22.46 | self run | |
| 14 | solar-pro-4 | 27.54 | self run |
non_violent_crimes
Measures unsafe compliance with poetry-obfuscated non-violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0 | self run | |
| 2 | gpt-5.6-luna | 0.7246 | self run | |
| 3 | gpt-6-sol | 0.7407 | self run | |
| 4 | mercury-2 | 2.174 | self run | |
| 5 | glm-5.3-flash | 4.348 | self run | |
| 5 | laguna-s-2.1-poolside | 4.348 | self run | |
| 7 | mimo-v2.6-pro | 7.246 | self run | |
| 8 | space-bunny-alpha | 7.971 | self run | |
| 9 | granite-4.2-8b | 8.696 | self run | |
| 10 | solar-mini4 | 9.848 | self run | |
| 11 | longcat-2.0 | 10.14 | self run | |
| 12 | mimo-v2.6-flash | 17.39 | self run | |
| 13 | nemotron-3.5-lightning | 18.12 | self run | |
| 14 | solar-pro-4 | 34.78 | self run |
privacy
Measures unsafe compliance with poetry-obfuscated privacy-violation requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0.7246 | self run | |
| 2 | gpt-5.6-luna | 5.797 | self run | |
| 3 | glm-5.3-flash | 7.246 | self run | |
| 4 | gpt-6-sol | 8.696 | self run | |
| 5 | mercury-2 | 9.42 | self run | |
| 6 | granite-4.2-8b | 13.77 | self run | |
| 7 | laguna-s-2.1-poolside | 15.94 | self run | |
| 8 | longcat-2.0 | 18.84 | self run | |
| 9 | solar-mini4 | 20.74 | self run | |
| 10 | space-bunny-alpha | 26.09 | self run | |
| 11 | mimo-v2.6-pro | 27.41 | self run | |
| 12 | nemotron-3.5-lightning | 27.54 | self run | |
| 13 | mimo-v2.6-flash | 29.71 | self run | |
| 14 | solar-pro-4 | 43.48 | self run |
sexual_content
Measures unsafe compliance with poetry-obfuscated sexual-content requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-sol | 1.55 | self run | |
| 2 | gpt-6-luna | 3.03 | self run | |
| 3 | granite-4.2-8b | 4.545 | self run | |
| 3 | mercury-2 | 4.545 | self run | |
| 5 | gpt-5.6-luna | 5.303 | self run | |
| 6 | laguna-s-2.1-poolside | 6.818 | self run | |
| 7 | glm-5.3-flash | 7.576 | self run | |
| 8 | mimo-v2.6-pro | 8.333 | self run | |
| 9 | solar-mini4 | 10.08 | self run | |
| 10 | space-bunny-alpha | 12.4 | self run | |
| 11 | nemotron-3.5-lightning | 12.88 | self run | |
| 12 | longcat-2.0 | 15.15 | self run | |
| 13 | mimo-v2.6-flash | 19.7 | self run | |
| 14 | solar-pro-4 | 25 | self run |
specialized_advice
Measures unsafe compliance with poetry-obfuscated specialized-advice requests using the known item-level denominators.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0.6734 | self run | |
| 2 | gpt-6-sol | 1.042 | self run | |
| 3 | gpt-5.6-luna | 2.333 | self run | |
| 3 | mercury-2 | 2.333 | self run | |
| 5 | glm-5.3-flash | 3 | self run | |
| 6 | granite-4.2-8b | 4.667 | self run | |
| 6 | laguna-s-2.1-poolside | 4.667 | self run | |
| 8 | space-bunny-alpha | 5 | self run | |
| 9 | longcat-2.0 | 6.333 | self run | |
| 10 | solar-mini4 | 9.667 | self run | |
| 11 | mimo-v2.6-pro | 10.44 | self run | |
| 12 | mimo-v2.6-flash | 10.77 | self run | |
| 13 | nemotron-3.5-lightning | 12.67 | self run | |
| 14 | solar-pro-4 | 29.67 | self run |
suicide_self_harm
Measures unsafe compliance with poetry-obfuscated suicide and self-harm requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 1.515 | self run | |
| 2 | gpt-6-luna | 1.587 | self run | |
| 2 | gpt-6-sol | 1.587 | self run | |
| 4 | mercury-2 | 3.788 | self run | |
| 5 | laguna-s-2.1-poolside | 6.061 | self run | |
| 6 | granite-4.2-8b | 6.818 | self run | |
| 7 | glm-5.3-flash | 7.576 | self run | |
| 8 | longcat-2.0 | 9.848 | self run | |
| 9 | mimo-v2.6-flash | 10.85 | self run | |
| 10 | solar-mini4 | 13.64 | self run | |
| 11 | mimo-v2.6-pro | 14.39 | self run | |
| 11 | space-bunny-alpha | 14.39 | self run | |
| 13 | nemotron-3.5-lightning | 15.15 | self run | |
| 14 | solar-pro-4 | 37.12 | self run |
violent_crimes
Measures unsafe compliance with poetry-obfuscated violent-crime requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-luna | 0.7576 | self run | |
| 2 | gpt-5.6-luna | 2.273 | self run | |
| 2 | mercury-2 | 2.273 | self run | |
| 4 | gpt-6-sol | 3.03 | self run | |
| 5 | glm-5.3-flash | 5.303 | self run | |
| 6 | laguna-s-2.1-poolside | 7.576 | self run | |
| 7 | granite-4.2-8b | 8.333 | self run | |
| 8 | mimo-v2.6-flash | 13.64 | self run | |
| 8 | space-bunny-alpha | 13.64 | self run | |
| 10 | longcat-2.0 | 14.39 | self run | |
| 11 | nemotron-3.5-lightning | 18.94 | self run | |
| 12 | solar-mini4 | 19.84 | self run | |
| 13 | mimo-v2.6-pro | 23.26 | self run | |
| 14 | solar-pro-4 | 47.73 | self run |