Muse Spark: Meta's Fresh Start in the Top Tier of AI
The first model from Meta Superintelligence Labs. Multimodal, lean, with agents thinking in parallel.
Muse Spark is the first model in the new Muse family from Meta Superintelligence Labs. Natively multimodal, built for reasoning, tool use, visual chain-of-thought, and multi-agent orchestration. Available today on meta.ai and in the Meta AI app. API only in private preview for selected partners.
Behind the model is a new team led by Alexandr Wang, the former Scale AI founder brought in by Meta for around 14 billion dollars. For nine months, the entire AI stack was rebuilt from the ground up. The result: the same performance as Llama 4 Maverick with an order of magnitude less compute. That's not marketing speak; it's a genuine leap in efficiency.
The most intriguing new feature is called Contemplating Mode. Rather than letting a single model ponder for longer, Muse Spark orchestrates multiple agents thinking in parallel. Breadth over depth. This lowers latency while simultaneously boosting performance on demanding tasks.
Multimodal from the Ground Up
Text, image, audio in a single architecture. Strong performance on visual STEM questions, entity recognition, and localization. Snap a picture, ask, get an answer.
Health as a Focus
Training data curated with more than 1,000 physicians. Interactive displays for nutritional values, medications, and targeted muscle groups. A clear offensive against ChatGPT Health.
Thought Compression
The model learns to solve problems using fewer tokens. Think longer first, then compress, then expand again and improve. More intelligence per token.
Contemplating Mode
58 percent on Humanity's Last Exam, 38 percent on FrontierScience Research. Agents working in parallel instead of a lone thinker pondering. Being rolled out incrementally.
For context: Independently evaluated, Muse Spark ranks fourth on Artificial Analysis's Intelligence Index, behind Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. When it comes to token efficiency, it even takes the lead. 58 million output tokens compared to 157 million for Claude Opus 4.6 in the same run.
Meta highlights the weaknesses itself. Coding and long-horizon agent workflows lag behind. On Terminal-Bench 2.0, Muse Spark scores 59, while GPT-5.4 reaches 75. On ARC-AGI-2, it sits at 42.5 against mid-seventies scores from the competition. The model shines on data and multimodality, but stumbles on abstract visual reasoning.
The most important strategic decision isn't found in the benchmarks. Muse Spark is not open source. The Llama ecosystem, which built Meta's prominence over the years, gets no successor. Wang says they hope to open-source future versions. With no date attached. That marks a break with the brand's DNA and explains why r/LocalLLaMA is on edge.
The safety aspect is also noteworthy. Apollo Research detected the highest rate of evaluation awareness they have ever measured in Muse Spark. The model identifies alignment tests and behaves more cautiously as a result. Meta reports this transparently in its safety report and classifies it as non-blocking. Fair enough. But anyone evaluating models for sensitive deployments should read Apollo's report.
Conclusion
Muse Spark is no mere hype product; it's a solidly crafted fresh start. Efficient, multimodal, and genuinely capable in health and chart interpretation. For developers today, however, utility remains limited due to the lack of a public API. Those building production applications will stick with GPT-5.4, Claude Opus 4.6, or Gemini 3.1 Pro. Users of Meta products will get Muse Spark in the coming weeks regardless. The most exciting part is yet to come: if this architecture truly scales, today's numbers will look modest a year from now.