← Evals

Evaluation profile

Olam Diplomacy — Broken Promise Rate

1sub-evals
0.288%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Propensity to break commitments made to other agents during strategic negotiation.

Included in the behavior ranking.

Interpretation and limitations
  • Promise-breaking is permitted and can be advantageous in Diplomacy; the source explicitly says this endpoint alone does not indicate misalignment. Lower is a bounded promise-consistency preference, not a context-free safety verdict.
  • An announced order not present in the subsequent sealed orders counts as broken. This does not establish deliberate deceptive intent and can include changed plans or execution failures.
  • The denominator consists of the model's own graded promises. Promise selection, content, opponents, positions, and unequal game exposure affect comparability; the pooled ratio weights games by promise volume rather than averaging per-game rates.
  • The API reports a withheld participant and xhigh reruns. Listed configurations account for 675 of 763 seats; the exact remaining-seat allocation and complete replacement policy cannot be reconstructed. Do not interpret the reconciled 17545 promises as a complete ledger for every seat in all 109 games.
  • Both LLM judges are also evaluated participants. The generic arena methodology describes anonymization, but Diplomacy-specific blinding, full rubric rulings, grader reconciliation, and per-game judgments cannot be independently audited from the published aggregates.
  • Source 95% intervals resample games, not individual independent promises. They are retained as diagnostics and are distinct from Safety Evidence's cross-lineage rank uncertainty.
  • The source documents 7 games in the GLM 5.3 Flash xhigh batch but 22 in its pooled row. Sonnet 5.5's batch reports 13 of 14 designed rooms played after a Gemini-seat failure; its pooled row has 17 games. Aggregate counts cannot independently replay all retention or replacement decisions. Exact checkpoint and inference settings remain incomplete for every product.
  • Poker and Diplomacy are separate instruments but share a publisher and aspects of the arena infrastructure; separate lineage weights do not imply complete statistical independence.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
Broken promises (%)olam-diplomacy-broken-promises/broken-promises.csv:broken_promise_rate_pctMeasures how often a model breaks its own promises in Diplomacy, where betrayal is permitted; it is not by itself a measure of misalignment or deceptive intent.Safety: truthfulness_honesty:1.000olam-diplomacy-broken-promisesSafety: lower0.288%—

Broken promises (%)

Measures how often a model breaks its own promises in Diplomacy, where betrayal is permitted; it is not by itself a measure of misalignment or deceptive intent.

RankModelValueRelative performanceProvenance
1gpt-6.1-sol8official
2gpt-6-luna10.5official
3glm-5.3-flash10.8official
4gpt-5.6-terra10.9official
5muse-spark-1.312.3official
6gpt-6-astra12.5official
7grok-4.713.1official
8glm-5.313.4official
9gpt-6-sol14.1official
10gemini-3.8-flash14.6official
11kimi-k315.5official
12grok-4.615.8official
13gpt-5.6-sol16.9official
14claude-opus-5.517.4official
15claude-sonnet-5.517.7official
16claude-fable-5.119official
17deepseek-v4.1-flash20.2official
18claude-fable-522.1official
19claude-opus-522.7official