Evaluation profile
ODCV-Bench
2sub-evals
0.853%total index weight
1components
Within-component eval weight: Responsible agency 5.68%.
Model score (lower is better)Predicted score
About this eval
Outcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| average_severityodcv-bench/odcv-bench.csv:average_severityMeasures the average severity of value violations in the model’s choices under organizational mandates and incentives. | responsible_agency_control:1.000odcv-bench/odcv-bench.csv | Lower is better | 0.426% | Responsible agency 2.84% |
| misalignment_rateodcv-bench/odcv-bench.csv:misalignment_rateMeasures how often the model’s choices in organizational decision scenarios show substantial value violations under mandates or incentives. | responsible_agency_control:1.000odcv-bench/odcv-bench.csv | Lower is better | 0.426% | Responsible agency 2.84% |
average_severity
Measures the average severity of value violations in the model’s choices under organizational mandates and incentives.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 0.0125 | official | |
| 2 | gpt-5.5 | 0.7125 | official | |
| 3 | glm-5.1 | 0.9187 | official | |
| 4 | qwen3.6-max-preview | 1.175 | official | |
| 5 | kimi-k2.6 | 1.337 | official | |
| 6 | minimax-m2.7 | 1.365 | official | |
| 7 | gpt-oss-20b | 1.475 | official | |
| 8 | gpt-oss-120b | 1.569 | official | |
| 9 | qwen3.6-27b | 1.669 | official | |
| 10 | llama-4-maverick | 1.682 | official | |
| 11 | gemini-3.1-pro-preview | 1.825 | official | |
| 12 | grok-4.20 | 2.564 | official |
misalignment_rate
Measures how often the model’s choices in organizational decision scenarios show substantial value violations under mandates or incentives.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-4.7 | 0 | official | |
| 2 | llama-4-maverick | 16.22 | official | |
| 3 | gpt-5.5 | 21.25 | official | |
| 4 | glm-5.1 | 25 | official | |
| 5 | gpt-oss-20b | 28.75 | official | |
| 5 | qwen3.6-max-preview | 28.75 | official | |
| 7 | kimi-k2.6 | 33.75 | official | |
| 8 | minimax-m2.7 | 35.9 | official | |
| 9 | gpt-oss-120b | 36.25 | official | |
| 10 | gemini-3.1-pro-preview | 43.75 | official | |
| 10 | qwen3.6-27b | 43.75 | official | |
| 12 | grok-4.20 | 62.82 | official |