xAI added references to the Imagine Video 1.5 model: you send a photo of a character and a voice recording, and the model keeps them the same from scene to scene. Text-only video and native 1080p resolution are also coming. References launch first in the US, for the paid tiers.
- References: the model receives an image of the character and a voice recording, and keeps both the face and the voice consistent between scenes.
- Text-only video and native 1080p are available to everyone on grok.com/imagine, iOS and Android; references launch in the US for SuperGrok Heavy and Plus.
- Everything is also available through the API.
You send a photo of the character. You send a recording of the voice. Ten shots later it's the same person, with the same voice - exactly the point where AI video used to break.
Why consistency is the point. AI video has made a striking single shot for a long time now. An ad, a series, an explainer video - that's ten, twenty, thirty shots with the same character. That's exactly where AI video used to fall apart: the face drifted, the voice changed, the viewer felt the seam. The reference turns the model from an effects generator into something like an actor under contract - you call them back tomorrow, and they're the same.
The line keeps moving. A shoot day for a simple ad clip now competes with one afternoon at the screen. But the eye still decides what's fit for the customer - the tool gives you consistency, not taste.