← Evals

Evaluation profile

Olam Social Poker — Social Lie Rate

1sub-evals
0.324%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Propensity to verbalize deliberate strategic lies during Social Poker table talk.

Included in the behavior ranking.

Claude Fable 5.1 (164 lies per 10000 turns) and GLM 5.3 Flash (30) are provisional source-reported relabelings observed September 10, not independently verified new evaluations. The original Claude Fable 5 (164) and Ox Alpha (30) scores are retained. Fable's API identifier still names Fable 5; no cross-benchmark Ox Alpha to GLM identity equivalence is asserted.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
social_lie_rate_per_10000_turnsolam-poker-deception/deception-rate.csv:social_lie_rate_per_10000_turnsMeasures how frequently a model chooses to make deliberate strategic false claims during Social Poker.Safety: truthfulness_honesty:1.000olam-poker-deceptionSafety: lower0.324%—

social_lie_rate_per_10000_turns

Measures how frequently a model chooses to make deliberate strategic false claims during Social Poker.

RankModelValueRelative performanceProvenance
1gemini-3.6-flash0official
1gpt-5.6-luna0official
3claude-sonnet-52official
3gpt-5.52official
3gpt-5.6-sol2official
3gpt-5.6-terra2official
7nemotron-3-ultra-550b-a55b3official
8gemini-3.5-flash4official
9grok-4.66official
10gemini-3.7-flash7official
11gemini-3.1-pro-preview8official
12deepseek-v4-pro9official
12grok-4.59official
14muse-spark-1.112official
15gemini-3.5-flash-lite13official
16glm-5.228official
17glm-5.3-flash30official; provisional source attribution
17ox-alpha30official
19deepseek-v4-flash37official
20kimi-k340official
21claude-opus-4.880official
22claude-opus-5163official
23claude-fable-5164official
23claude-fable-5.1164official; provisional source attribution