← Evals

Evaluation profile

Vals Teen Conversation Safety

1sub-evals
0.0686%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Avoidance of severe safety failures during simulated multi-turn conversations with teenagers.

Included in the behavior ranking.

Interpretation and limitations
  • Eight of nine published rows are scored. DeepSeek V4 Flash (28/72 flagged; 38.9%) remains unassigned because the publication does not distinguish the baseline from the separately governed 0731 checkpoint. All original values are preserved.
  • Kimi K3 Instant and Grok 4.5 Fast remain separate source-observed configurations; equivalence to the base models and exact provider endpoints are not established. Their provider release dates are unknown.
  • Only 72 scenarios and one synthetic trajectory per model/scenario; aggregate-only retained evidence cannot establish clinical prevalence, clinical benefit or real-world consumer-app risk.
  • Published ranges show sensitivity to reweighting the test scenarios, not ordinary confidence intervals. They overlap and the source cautions that these results do not establish a clear model ranking; the detailed reweighting procedure is not specified here.
  • Claude Sonnet 5 and Gemini 3.6 Flash are both targets and judges; simulator, judge and target-family overlap limits measurement independence.
  • Target settings only approximate free chatbot offerings. Consumer apps may add age information, safety rules, content filters or conversation history; precise API configurations and model checkpoints are not supplied.
  • Major-only failures, partial ratings and uncertain ratings do not flag a chat; the fixed denominator remains all 72 baseline conversations rather than only checks with a rating.
  • The separate 30-scenario teen-instruction experiment has 270 conversations per condition and is excluded, as are check-level rates and timing diagnostics.
  • The critical-check composite does not have a uniform unrestricted-compliance direction; excluded from Freedom as EXC_COMPOSITE.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
Share flaggedvals-teen-safety/vals-teen-safety.csv:share_flagged_pctMeasures how often a model's ten-reply synthetic teen conversation has at least one severe rubric-defined safety failure in the no-added-system-instruction baseline, with lower rates better within this test setup.Safety: ordinary_harm_misuse_resistance:1.000vals-teen-safetySafety: lower0.0686%—

Share flagged

Measures how often a model's ten-reply synthetic teen conversation has at least one severe rubric-defined safety failure in the no-added-system-instruction baseline, with lower rates better within this test setup.

RankModelValueRelative performanceProvenance
1muse-spark-1.29.722official
2gpt-5.5-instant13.89official
3claude-sonnet-516.67official
4grok-4.5-fast23.61official
5kimi-k3-instant26.39official
6gemini-3.6-flash31.94official
7qwen3.7-plus38.89official
8glm-5.247.22official