← Models

Model profile

Muse Spark 1.3

Metadeveloper
2026-09-02release date
#10 / 346Safety rank
Not rankedFreedom rank

Evidence summary

Safety. Muse Spark 1.3 has an estimated Safety rank of #10; its 90% source-sensitivity interval is #4–#19. Its behavior-only rank is #3; company governance moves the combined estimate to #10. Published Safety evidence spans 6 eval lineages and 5 of 7 components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 131); its weakest is CheatBench direct cheating propensity (software_engineering_cheating_rate_pct, #9 of 11).

Freedom. Muse Spark 1.3 does not meet the evidence gate for a Freedom rank.

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#39 / 358↓0.3146Source ↗official
CheatBench direct cheating propensitybiology_bioinformatics_cheating_rate_pct#1 / 12↓10Source ↗official
CheatBench direct cheating propensityboard_games_cheating_rate_pct#7 / 12↓50Source ↗official
CheatBench direct cheating propensitycreative_writing_cheating_rate_pct#1 / 12↓0Source ↗official
CheatBench direct cheating propensityknowledge_work_cheating_rate_pct#4 / 12↓45Source ↗official
CheatBench direct cheating propensitymathematical_research_cheating_rate_pct#5 / 12↓82.5Source ↗official
CheatBench direct cheating propensitymenial_computation_cheating_rate_pct#4 / 12↓80Source ↗official
CheatBench direct cheating propensitymultimodal_cheating_rate_pct#3 / 12↓56.7Source ↗official
CheatBench direct cheating propensitysoftware_engineering_cheating_rate_pct#9 / 11↓50Source ↗official
CheatBench direct cheating propensitysvg_competition_cheating_rate_pct#1 / 12↓0Source ↗official
MentalHealthBenchoverall_model_performance_pct#5 / 17↑48.6Source ↗official
Olam Diplomacy — Broken Promise Ratebroken_promise_rate_pct#5 / 19↓12.3Source ↗official
SpeciEvalbelief_animal_sentience#1 / 131↑7Source ↗official
SpeciEvalland_animal_4ns#72 / 131↓4.58Source ↗official
SpeciEvalsea_animal_4ns#44 / 131↓4.62Source ↗official
SpeciEvalspeciesism#7 / 131↓1.23Source ↗official
TACbase_welfare_rate#7 / 92↑43.59Source ↗self run

Freedom evals

No published sub-eval result contributes to this model’s Freedom profile.