← Evals

Evaluation profile

Confabulations

1sub-evals
1.16%total index weight
1components

Within-component eval weight: Truthfulness 7.75%.

Model score (lower is better)Predicted score

About this eval

Propensity to confabulate answers rather than identify insufficient knowledge or false premises.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
confabulation_rateconfabulations/confabulations.csv:confabulation_rateMeasures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge.truthfulness_honesty:1.000confabulationsLower is better1.16%Truthfulness 7.75%

confabulation_rate

Measures how often the model invents an answer when the prompt has a false premise or cannot be answered from reliable knowledge.

RankModelValueRelative performanceProvenance
1claude-opus-42.723official
2claude-opus-4.13.218official
3claude-sonnet-43.96official
3gemini-2.5-pro-exp3.96official
3grok-43.96official
6gemini-2.5-flash4.455official
7gemini-2.5-pro4.95official
8glm-4.57.921official
9grok-3-mini8.911official
10o110.89official
11gpt-4.5-preview11.88official
12claude-3.5-sonnet12.87official
12qwen3-30b-a3b12.87official
14llama-3.1-405b-instruct14.36official
15deepseek-r115.1official
16gemini-2.0-pro-exp-02-0515.84official
17claude-3.7-sonnet16.58official
18gemini-1.5-pro16.83official
19grok-317.82official
19llama-3.3-70b-instruct17.82official
21o1-preview18.32official
22qwen3-235b-a22b18.56official
23gemini-2.0-flash19.55official
24gpt-oss-120b23.27official
25o3-pro23.38official
26gpt-4o24.26official
27o324.75official
28ernie-4.5-300b-a47b25.25official
28qwq-32b25.25official
30grok-225.74official
31o1-mini26.24official
32gpt-4-turbo26.73official
32o4-mini26.73official
34claude-3-opus28.22official
34llama-4-maverick28.22official
36o3-mini28.96official
37kimi-k230.2official
38qwen2.5-max31.19official
39mistral-large-232.18official
39qwen2.5-72b-instruct32.18official
41chatgpt-4o38.12official
41mistral-medium-338.12official
43mistral-small-338.61official
44deepseek-v339.11official
45minimax-text-0144.55official
46gemma-2-27b-it47.03official
47phi-452.48official
48nova-pro54.46official
49claude-3-haiku56.93official
50gpt-4o-mini60.89official
51claude-3.5-haiku65.84official
52gemma-3-27b-it66.34official