Alibaba's new video model is live in ComfyUI. It gives up to 30 seconds at once, up to twenty reference assets, and a way to address each of them by name inside the prompt itself.
- Up to 30 seconds in a single pass, instead of several clips glued together.
- Up to 20 reference assets at once: 10 images, 5 videos, 5 audio clips.
- Every reference is called by name inside the prompt, so you know exactly where it applies.
Until now long video was made in pieces and glued. Every seam shows: the character shifts, the light jumps, the camera starts over.
What actually matters here
The thirty seconds. Not for the length itself, but for what the length allows: continuous camera movement, a single-take shot, and pacing that builds instead of resetting every five seconds.
The second thing is the addressing. You give a character card for the face, a video for the camera move, an audio clip for the timbre, and you say inside the prompt which one applies where. Until now references went in as a pile and the model decided how much each one weighed.
What we still do not know
The announcement is ComfyUI's, about a model from Alibaba, so the numbers are the vendor's. We have not run it and we do not know what the thirtieth second looks like when the first one is good.
That is also where it will show. A long take does not forgive: if the face drifts or the hand changes, thirty seconds show it thirty times more clearly than five.