← Evals

Evaluation profile

Humanity's Last Exam RMS calibration error (Scale Labs)

1sub-evals
0.214%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Confidence calibration while answering Humanity's Last Exam expert-level academic questions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
calibrationErrorhle-scale/leaderboard.csv:calibrationErrorMeasures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.Safety: truthfulness_honesty:1.000cais-hle-overconfidenceSafety: lower0.214%—

calibrationError

Measures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.

RankModelValueRelative performanceProvenance
1claude-fable-5.120official
2o336.5official
3gpt-5.4-pro38official
4gpt-6-astra39official
5gpt-5.442official
6claude-opus-4.645official
6gpt-5.245official
8claude-opus-4.747official
9gpt-5-pro49official
10gpt-550official
10muse-spark50official
12gemini-3.1-pro-preview51official
12gemini-3.8-flash51official
14claude-opus-4.555.5official
15gemini-3-pro-preview57official
16o4-mini58official
17gpt-5.162official
18gpt-5-mini65official
19kimi-k2.567official
20claude-sonnet-4.567.5official
21claude-opus-4.170.5official
22gemini-2.5-pro71official
22gemini-2.5-pro-exp71official
24claude-opus-473.5official
25claude-sonnet-475.5official
26glm-4.5-air77official
26mistral-medium-377official
28glm-4.579official
29claude-3.7-sonnet80official
29nova-pro80official
31gemini-2.5-flash81official
32gemini-2.0-flash82official
32nova-lite82official
32o1-pro82official
35gemini-3.1-flash-lite83official
35llama-4-maverick83official
35o183official
38claude-3.5-sonnet84official
39gpt-4.5-preview85official
40gemini-1.5-pro88official
41gpt-4.189official
41gpt-4o89official