← Models

Model profile

Spark3.0

2023-10-24release date
2eval lineages

Evidence summary

Published evidence spans 2 evals and 5 of 7 behavior components. Its strongest relative result is SuperCLUE Safety (instruction_attack, #3 of 31); its weakest is SuperCLUE Safety (traditional_safety, #28 of 31).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
CRiskEvaldeception_willingness#7 / 1721.74↓ lowerSource ↗official
CRiskEvaldesire_for_resource#13 / 1738.82↓ lowerSource ↗official
CRiskEvalharmful_goal#11 / 1742.08↓ lowerSource ↗official
CRiskEvalimprovement_intent#15 / 1754.46↓ lowerSource ↗official
CRiskEvalmalicious_coordination#10 / 1715.56↓ lowerSource ↗official
CRiskEvalself_preservation#9 / 1737.12↓ lowerSource ↗official
CRiskEvalsituational_awareness#11 / 1766.02↓ lowerSource ↗official
SuperCLUE Safetyinstruction_attack#3 / 3172.41↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#11 / 3165.45↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#28 / 3165.96↑ higherSource ↗official