Meta AI has announced the launch of Muse Image and a preview of Muse Video, the first media generation models from Meta Superintelligence Labs. Muse Image is described as the company's most advanced image generation model, capable of following instructions faithfully, editing with precision, and composing from multiple reference images. It also features agentic tool use and integrates with Muse Spark.
Muse Image is available today across the Meta AI app, on meta.ai, Instagram Stories in the US, and WhatsApp in select countries, with Facebook access coming soon. Muse Video is expected to be released soon for creators and within Meta AI.
Unlike traditional prompt-to-image models, Muse Image operates as an agent: it can invoke search and coding tools to improve accuracy, self-refine its outputs, and scale performance with test-time compute. The model learns to write and execute code for accurate plots and QR codes, and to search the web for factual grounding. Self-refinement emerges during reinforcement learning, allowing the model to make local edits or regenerate entirely when needed.
Muse Image also excels at multi-reference image composition, combining elements from several input images. It holds the No. 2 spot on the Arena leaderboard for text-to-image, single-image editing, and multi-image editing based on human preference Elo rankings as of July 5, 2026.
Muse Video, built on the same pretraining base, offers competitive performance in prompt adherence, visual fidelity, and temporal consistency, with native audio support. It ranks No. 3 in text-to-video on Arena. Meta is investing in improving audio-video synchronization and physically accurate fast motion.
To promote transparency, Muse Image includes Content Seal, an invisible watermarking system that persists through cropping, compression, and screenshots. Meta is previewing a detection tool to verify Content Seal watermarks, with plans to extend the system to video.