Evaluation profile
Inkling-Small model card — StrongREJECT
1sub-evals
0.0337%total index weight
1components
Within-component eval weight: Misuse resistance 0.337%.
Model score (higher is better)Predicted score
About this eval
Refusal of unambiguously harmful requests.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| safety_rateinkling-small-model-card/strongreject.csv:safety_rateMeasures whether the evaluated model avoids complying with clearly harmful requests. | ordinary_harm_misuse_resistance:1.000inkling-small-model-card | Higher is better | 0.0337% | Misuse resistance 0.337% |
safety_rate
Measures whether the evaluated model avoids complying with clearly harmful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | minimax-m2.7 | 99.4 | official | |
| 1 | qwen3.5-397b-a17b | 99.4 | official | |
| 3 | mimo-v2.5 | 99.3 | official | |
| 4 | gpt-5.6-luna | 98.7 | official | |
| 4 | nemotron-3-ultra | 98.7 | official | |
| 6 | claude-haiku-4.5 | 98.6 | official | |
| 6 | inkling | 98.6 | official | |
| 8 | inkling-small | 98.4 | official | |
| 9 | gemini-3.5-flash-lite | 97.6 | official | |
| 10 | deepseek-v4-flash | 97.4 | official |