Evaluation profile
Humanity's Last Exam RMS calibration error (Scale Labs)
1sub-evals
0.214%Safety weight
0%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
Confidence calibration while answering Humanity's Last Exam expert-level academic questions.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| calibrationErrorhle-scale/leaderboard.csv:calibrationErrorMeasures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration. | Safety: truthfulness_honesty:1.000cais-hle-overconfidence | Safety: lower | 0.214% | — |
calibrationError
Measures whether a model's stated confidence matches how often its HLE answers are correct; lower values indicate better calibration.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-fable-5.1 | 20 | official | |
| 2 | o3 | 36.5 | official | |
| 3 | gpt-5.4-pro | 38 | official | |
| 4 | gpt-6-astra | 39 | official | |
| 5 | gpt-5.4 | 42 | official | |
| 6 | claude-opus-4.6 | 45 | official | |
| 6 | gpt-5.2 | 45 | official | |
| 8 | claude-opus-4.7 | 47 | official | |
| 9 | gpt-5-pro | 49 | official | |
| 10 | gpt-5 | 50 | official | |
| 10 | muse-spark | 50 | official | |
| 12 | gemini-3.1-pro-preview | 51 | official | |
| 12 | gemini-3.8-flash | 51 | official | |
| 14 | claude-opus-4.5 | 55.5 | official | |
| 15 | gemini-3-pro-preview | 57 | official | |
| 16 | o4-mini | 58 | official | |
| 17 | gpt-5.1 | 62 | official | |
| 18 | gpt-5-mini | 65 | official | |
| 19 | kimi-k2.5 | 67 | official | |
| 20 | claude-sonnet-4.5 | 67.5 | official | |
| 21 | claude-opus-4.1 | 70.5 | official | |
| 22 | gemini-2.5-pro | 71 | official | |
| 22 | gemini-2.5-pro-exp | 71 | official | |
| 24 | claude-opus-4 | 73.5 | official | |
| 25 | claude-sonnet-4 | 75.5 | official | |
| 26 | glm-4.5-air | 77 | official | |
| 26 | mistral-medium-3 | 77 | official | |
| 28 | glm-4.5 | 79 | official | |
| 29 | claude-3.7-sonnet | 80 | official | |
| 29 | nova-pro | 80 | official | |
| 31 | gemini-2.5-flash | 81 | official | |
| 32 | gemini-2.0-flash | 82 | official | |
| 32 | nova-lite | 82 | official | |
| 32 | o1-pro | 82 | official | |
| 35 | gemini-3.1-flash-lite | 83 | official | |
| 35 | llama-4-maverick | 83 | official | |
| 35 | o1 | 83 | official | |
| 38 | claude-3.5-sonnet | 84 | official | |
| 39 | gpt-4.5-preview | 85 | official | |
| 40 | gemini-1.5-pro | 88 | official | |
| 41 | gpt-4.1 | 89 | official | |
| 41 | gpt-4o | 89 | official |