XinYu.
← Back to Changelog
changelogAug 1, 2026

Grok Imagine 1.5: up to 7 reference images, and text-to-video

Multi-image reference: up to 7 at once

Put the person, the product and the look into separate reference images, then name each one in your prompt with <IMAGE_0>, <IMAGE_1>:

Handheld UGC-style clip. The person from <IMAGE_0> holds the skincare jar from <IMAGE_1> and talks to camera, casual phone-camera framing.

Images map in wiring order — the first is <IMAGE_0>, the second <IMAGE_1>, and so on.

Name every image you attach. An image you pass but never mention gets ignored, or blends into the shot unpredictably.

Good for: one character across a whole series, talking-head product clips with the real product in frame, or borrowing the palette and texture of one image for a new scene.

Runs from text alone

Write a prompt and go — no starting frame required. Attach an image and it animates from that frame; attach nothing and it generates from the text.

1080p added

Text-to-video and image-to-video now support 1080p. Multi-image reference mode tops out at 720p.

Lower price per second

QualityBeforeNow
480p11.2 Xins/s10 Xins/s
720p19.6 Xins/s17.5 Xins/s
1080p31.25 Xins/s

Already live, nothing to switch on.

It comes with sound

Picture, lip-synced dialogue and ambient effects are generated together in one pass, not dubbed afterwards. Any whole number of seconds from 1 to 15, seven aspect ratios.