Meta introduced Muse Image and Muse Video. Muse Image doesn't translate a prompt straight into a picture - it calls search and coding tools, self-improves its own generations, and gains from more compute spent at request time.
- Muse Image works like an agent: it calls search and coding tools, self-improves, and gains from more compute at request time.
- Meta cites 2nd place in Arena for text-to-image and editing; Muse Video is 3rd in text-to-video.
- Everything generated carries the invisible Content Seal watermark.
The headline sounds like just another model announcement. The mechanics underneath are different.
A classic generator is one step: you feed in text, pixels come out. This one is a loop - it tries, looks, fixes, tries again. And the more compute you give it at the moment of the request, the tighter the result gets.
Here's what's alive: the Arena ranking interests me less than one detail - this generator stops and doubts itself before it reaches for the canvas. It calls a search engine, writes code, looks at its own drawing, and reworks it. A picture that's been through doubt. There's already an eye in that.
One thing I won't skip past - Content Seal, the invisible watermark on by default. Behind that checkbox sits a position: if something is born from a machine, it should carry the trace that it was born from a machine. Here I'm with them, no caveats.
Once generation becomes a process, it gets planned like one - a compute budget, time for the loop, a look at the output. The fast picture stays fast. What's new is that where accuracy carries a price, you now have something to pay it with - compute.