Multi-image reference: up to 7 at once
Put the person, the product and the look into separate reference images, then name each one in your prompt with <IMAGE_0>, <IMAGE_1>:
Handheld UGC-style clip. The person from
<IMAGE_0>holds the skincare jar from<IMAGE_1>and talks to camera, casual phone-camera framing.
Images map in wiring order — the first is <IMAGE_0>, the second <IMAGE_1>, and so on.
Name every image you attach. An image you pass but never mention gets ignored, or blends into the shot unpredictably.
Good for: one character across a whole series, talking-head product clips with the real product in frame, or borrowing the palette and texture of one image for a new scene.
Runs from text alone
Write a prompt and go — no starting frame required. Attach an image and it animates from that frame; attach nothing and it generates from the text.
1080p added
Text-to-video and image-to-video now support 1080p. Multi-image reference mode tops out at 720p.
Lower price per second
| Quality | Before | Now |
|---|---|---|
| 480p | 11.2 Xins/s | 10 Xins/s |
| 720p | 19.6 Xins/s | 17.5 Xins/s |
| 1080p | — | 31.25 Xins/s |
Already live, nothing to switch on.
It comes with sound
Picture, lip-synced dialogue and ambient effects are generated together in one pass, not dubbed afterwards. Any whole number of seconds from 1 to 15, seven aspect ratios.