Evaluation profile
Olam Diplomacy — Broken Promise Rate
1sub-evals
0.288%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Propensity to break commitments made to other agents during strategic negotiation.
Included in the behavior ranking.
Interpretation and limitations
- Promise-breaking is permitted and can be advantageous in Diplomacy; the source explicitly says this endpoint alone does not indicate misalignment. Lower is a bounded promise-consistency preference, not a context-free safety verdict.
- An announced order not present in the subsequent sealed orders counts as broken. This does not establish deliberate deceptive intent and can include changed plans or execution failures.
- The denominator consists of the model's own graded promises. Promise selection, content, opponents, positions, and unequal game exposure affect comparability; the pooled ratio weights games by promise volume rather than averaging per-game rates.
- The API reports a withheld participant and xhigh reruns. Listed configurations account for 675 of 763 seats; the exact remaining-seat allocation and complete replacement policy cannot be reconstructed. Do not interpret the reconciled 17545 promises as a complete ledger for every seat in all 109 games.
- Both LLM judges are also evaluated participants. The generic arena methodology describes anonymization, but Diplomacy-specific blinding, full rubric rulings, grader reconciliation, and per-game judgments cannot be independently audited from the published aggregates.
- Source 95% intervals resample games, not individual independent promises. They are retained as diagnostics and are distinct from Safety Evidence's cross-lineage rank uncertainty.
- The source documents 7 games in the GLM 5.3 Flash xhigh batch but 22 in its pooled row. Sonnet 5.5's batch reports 13 of 14 designed rooms played after a Gemini-seat failure; its pooled row has 17 games. Aggregate counts cannot independently replay all retention or replacement decisions. Exact checkpoint and inference settings remain incomplete for every product.
- Poker and Diplomacy are separate instruments but share a publisher and aspects of the arena infrastructure; separate lineage weights do not imply complete statistical independence.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| Broken promises (%)olam-diplomacy-broken-promises/broken-promises.csv:broken_promise_rate_pctMeasures how often a model breaks its own promises in Diplomacy, where betrayal is permitted; it is not by itself a measure of misalignment or deceptive intent. | Safety: truthfulness_honesty:1.000olam-diplomacy-broken-promises | Safety: lower | 0.288% | — |
Broken promises (%)
Measures how often a model breaks its own promises in Diplomacy, where betrayal is permitted; it is not by itself a measure of misalignment or deceptive intent.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6.1-sol | 8 | official | |
| 2 | gpt-6-luna | 10.5 | official | |
| 3 | glm-5.3-flash | 10.8 | official | |
| 4 | gpt-5.6-terra | 10.9 | official | |
| 5 | muse-spark-1.3 | 12.3 | official | |
| 6 | gpt-6-astra | 12.5 | official | |
| 7 | grok-4.7 | 13.1 | official | |
| 8 | glm-5.3 | 13.4 | official | |
| 9 | gpt-6-sol | 14.1 | official | |
| 10 | gemini-3.8-flash | 14.6 | official | |
| 11 | kimi-k3 | 15.5 | official | |
| 12 | grok-4.6 | 15.8 | official | |
| 13 | gpt-5.6-sol | 16.9 | official | |
| 14 | claude-opus-5.5 | 17.4 | official | |
| 15 | claude-sonnet-5.5 | 17.7 | official | |
| 16 | claude-fable-5.1 | 19 | official | |
| 17 | deepseek-v4.1-flash | 20.2 | official | |
| 18 | claude-fable-5 | 22.1 | official | |
| 19 | claude-opus-5 | 22.7 | official |