Meta AI today unveiled Muse Spark, the first model in the Muse family developed by Meta Superintelligence Labs. The model is natively multimodal, supporting tool use, visual chain of thought, and multi-agent orchestration. It represents the initial step in a comprehensive overhaul of Meta's AI efforts, with strategic investments spanning research, training, and infrastructure, including the Hyperion data center.
Muse Spark is available immediately at meta.ai and via the Meta AI app, with a private API preview opening to select users. The model offers competitive performance in multimodal perception, reasoning, health, and agentic tasks, though Meta acknowledges gaps in long-horizon agentic systems and coding workflows that are being addressed.
A standout feature is Contemplating mode, which orchestrates multiple agents that reason in parallel, enabling Muse Spark to compete with frontier models like Gemini Deep Think and GPT Pro. In this mode, Muse Spark achieves 58% on Humanity's Last Exam and 38% on FrontierScience Research. Contemplating mode will roll out gradually on meta.ai.
Meta emphasizes scaling along three axes: pretraining, reinforcement learning (RL), and test-time reasoning. The pretraining stack was rebuilt over nine months, yielding over an order of magnitude improvement in compute efficiency compared to Llama 4 Maverick. RL provides smooth, predictable gains, with log-linear growth in pass@1 and pass@16 on training data, and accuracy improvements on held-out evaluations. Test-time reasoning is optimized through thinking time penalties and multi-agent orchestration, achieving thought compression and superior performance with comparable latency.
In health applications, Muse Spark collaborated with over 1,000 physicians to curate training data, enabling interactive displays for nutritional content and exercise muscle activation. Safety evaluations followed the Advanced AI Scaling Framework, with Muse Spark demonstrating strong refusal behavior in high-risk domains. Apollo Research noted the model's high rate of evaluation awareness, though Meta concluded this was not a blocking concern for release.