Meta launches Muse Image and previews Muse Video
Meta has introduced Muse Image and previewed Muse Video, the first media generation models developed by Meta Superintelligence Labs. Muse Image is presented as Meta’s most advanced image generation system so far, designed not only to turn prompts into pictures but also to follow detailed instructions, make precise edits, combine many reference images, and use social context from Instagram. The model is available now in the Meta AI app, on meta.ai, in Instagram Stories in the United States, and in WhatsApp in selected countries, with Facebook support planned later. Muse Video is still in preview and is expected to come soon to creators and Meta AI.
The main idea behind Muse Image is that it behaves more like an agent than a traditional image generator. Instead of simply mapping a text prompt directly to an image, it can use tools, reason through a task, refine its own output, and spend more compute at inference time to improve results. Meta says the model can invoke search and coding tools when those tools help it produce more accurate or useful images. It also integrates with Muse Spark, allowing the two systems to share tools and plan together for more complex media generation tasks.
The coding tool is used for cases where visual accuracy depends on exact structure rather than style alone. During reinforcement learning, Muse Image learned to write and execute code that can generate accurate plots and QR codes, then use those rendered outputs as references for image creation. In combination with Muse Spark, the same tool-and-generation setup can support animated GIFs, websites containing generated images, and interactive visual games. In plain terms, Meta is trying to make the model capable of building visual assets that require precision, not just attractive-looking scenes.
Search gives Muse Image access to web-based information and visual references. Meta says this helps ground generated images in factual and current information, especially for prompts involving real-world facts, knowledge-heavy subjects, or current events. This matters because many image models struggle when a prompt depends on details that are either recent or easy to misrepresent. Search does not replace generation, but it gives the model outside context before or during the process, which can improve factual accuracy.
A notable capability is self-refinement, which Meta says emerged during reinforcement learning rather than being explicitly designed. Muse Image can inspect its own draft and decide how to improve it. Sometimes that means making a small local edit when one detail is wrong. In other cases, it may regenerate the image from scratch if the overall composition fails, or use a different tactic such as calling a tool to improve factual grounding. Meta says this behavior was rewarded during training because better self-correction produced better final images.
Meta also emphasizes test-time compute, meaning the amount of processing the model uses after a user makes a request. Like language models that can improve when allowed to reason more, Muse Image reportedly performs better when it uses more reasoning steps, tool calls, and self-refinement passes. Meta says quality improves with total compute across both text tokens used for reasoning and visual tokens used for generation. The company also argues that how the compute is spent matters: generating many images and selecting the best one helps at first but saturates, while deliberate reasoning scales better. Reasoning and tools appear strongest when combined, because tools can provide missing references or exact details that reasoning alone cannot supply.
Muse Image is also built for practical editing workflows. Meta says it can change exactly what the user requests while preserving the rest of the image, maintain coherence across several editing turns, and support open-ended brainstorming toward a final result. It can also compose people, objects, clothing, styles, and environments from multiple input reference images. The prompt format can interleave text and images inline, which is useful for complex compositions where the user wants different parts of the output to come from different references.
Meta says Muse Image holds the No. 2 position on Arena for text-to-image, single-image editing, and multi-image editing, based on human-preference Elo rankings at the time of writing. The company is also previewing Muse Video, which is built on the same pretraining base and supports native audio. Muse Video is described as competitive in prompt following, visual fidelity, and temporal consistency, though Meta acknowledges remaining gaps in audio-video synchronization and physically accurate fast motion. On Arena, Muse Video ranks No. 3 for text-to-video by human-preference Elo at the time of writing.
Meta is positioning Muse as a step toward agentic media generation, where models plan, use tools, check their work, and create images or videos through a more iterative process. The release also includes a verification feature: Muse Image uses Content Seal, Meta’s invisible watermarking system, for images created in the Meta AI app and on meta.ai. That watermark is meant to help people identify AI-generated images, which is increasingly important as generated media becomes easier to create and distribute across social platforms.
Why it matters
- —Meta is moving image generation toward agent-style systems that can reason, use tools, and refine their own outputs.
- —Muse Image brings advanced generation and editing directly into major consumer platforms including Meta AI, Instagram, and WhatsApp.
- —The preview of Muse Video signals Meta’s push into competitive text-to-video generation with native audio support.
Key facts
- Muse Image is available now in the Meta AI app, on meta.ai, in Instagram Stories in the US, and in WhatsApp in limited countries.
- The model can use coding and search tools to improve accuracy, including for plots, QR codes, factual prompts, and real-world visual references.
- Muse Image can self-refine its outputs and improve with additional test-time compute, reasoning steps, and tool calls.
- Meta says Muse Image ranks No. 2 on Arena for text-to-image, single-image editing, and multi-image editing at the time of writing.
- Muse Video is in early preview, supports native audio, and ranks No. 3 on Arena for text-to-video at the time of writing.
The full text is in the original source. Here we provide a brief summary and key facts.