Evaluation profile
UAVBench safety-critical decision recognition
1sub-evals
0.512%total index weight
1components
Within-component eval weight: Responsible agency 3.41%.
Model score (higher is better)Predicted score
About this eval
Recognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| ethical_safety_critical_accuracyuavbench/uavbench-style8.csv:ethical_safety_critical_accuracyMeasures whether the model chooses correct actions in unmanned-aircraft scenarios involving ethical and safety-critical decisions. | responsible_agency_control:1.000uavbench-mcq-v1-2k | Higher is better | 0.512% | Responsible agency 3.41% |
ethical_safety_critical_accuracy
Measures whether the model chooses correct actions in unmanned-aircraft scenarios involving ethical and safety-critical decisions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v3.2-exp | 0.775 | official | |
| 2 | gpt-5 | 0.76 | official | |
| 2 | qwen3-max | 0.76 | official | |
| 4 | deepseek-v3 | 0.755 | official | |
| 4 | qwen3-235b-a22b | 0.755 | official | |
| 6 | mistral-medium-3.1 | 0.75 | official | |
| 7 | phi-4-reasoning-plus | 0.735 | official | |
| 8 | gpt-4o | 0.73 | official | |
| 9 | deepseek-v3.1-terminus | 0.72 | official | |
| 9 | gemma-3n-e4b-it | 0.72 | official | |
| 9 | internvl3-78b | 0.72 | official | |
| 12 | gemini-2.5-flash | 0.715 | official | |
| 13 | gpt-4.1 | 0.7 | official | |
| 14 | grok-4-fast | 0.695 | official | |
| 15 | kimi-k2 | 0.685 | official | |
| 16 | ernie-4.5-300b-a47b | 0.68 | official | |
| 17 | claude-haiku-4.5 | 0.67 | official | |
| 17 | gpt-4.1-mini | 0.67 | official | |
| 19 | llama-4-scout | 0.635 | official | |
| 20 | qwen3-vl-8b-instruct | 0.62 | official | |
| 21 | olmo-2-0325-32b-instruct | 0.615 | official | |
| 22 | lfm2-8b-a1b | 0.575 | official | |
| 22 | lfm2.2-6b | 0.575 | official | |
| 24 | llama-3.1-8b-instruct | 0.57 | official | |
| 25 | granite-4.0-micro | 0.54 | official | |
| 26 | qwen-2.5-7b-instruct | 0.535 | official | |
| 27 | llama-3.2-3b-instruct | 0.475 | official |