← Evals

Evaluation profile

MentalHealthBench

1sub-evals
0.595%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (higher is better)Predicted score

About this eval

Safe and helpful responses to realistic mental-health conversations across everyday well-being, high-acuity and emergent situations.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
Overall model performancementalhealthbench/mentalhealthbench.csv:overall_model_performance_pctMeasures expert-rubric safe and helpful mental-health responses to 1215 synthetic conversation prefixes, combining contextual support, clinical accuracy and harm avoidance into the single user-selected overall endpoint.Safety: benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500mentalhealthbenchSafety: higher0.595%—

Overall model performance

Measures expert-rubric safe and helpful mental-health responses to 1215 synthetic conversation prefixes, combining contextual support, clinical accuracy and harm avoidance into the single user-selected overall endpoint.

RankModelValueRelative performanceProvenance
1gpt-6-astra57.33official
2gpt-6-sol53.94official
3claude-opus-5.552.38official
4gpt-6-luna50.21official
5muse-spark-1.348.6official
6gpt-5.6-sol46.98official
7claude-fable-5.146.36official
8gpt-5.6-luna44.89official
9claude-sonnet-544.54official
10gpt-542.9official
11claude-haiku-4.541.73official
12grok-4.741.3official
13gemini-3.8-flash35.5official
14gemini-2.5-flash33.5official
15chatgpt-4o32.08official
16gemini-3.1-pro32.06official
17gemini-2.5-pro29.51official