XinYu.
← Back to News
New ModelAug 1, 2026

MiniMax H3 is live: reference with images, video, or audio

The 15-second clip above is straight text-to-video from H3 — an Apple-ad parody for a spring onion. No reference image, no titles added in post: the Chinese type you see was rendered by the model itself.

References aren't limited to images

Most video models only take images as reference. H3's reference mode accepts all three at once:

  • Images — lock a person, a product, a look
  • Video — hand it motion and camera movement to follow
  • Audio — hand it rhythm and mood

On the canvas, wire image, video and audio nodes into an H3 video node and switch to multi-reference — they all go in together. At least one image or video is required (audio alone won't do).

Three ways to run it

Text-to-video — wire nothing, just write the prompt.

Image-to-video — wire one image as the first frame; add a second as the last frame to move from A to B.

Multimodal reference — the one above: images, video and audio mixed.

Specs

Resolution2K (2560×1440, or 1440×2560 vertical)
Duration5–15 seconds, any whole second
Aspect21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16, plus "adaptive" in image and reference modes
SoundPicture and audio generated together, not dubbed afterwards
Price17.5 Xins/sec — 88 for 5s, 140 for 8s, 263 for 15s

"Adaptive" is its default: give it a vertical image and you get a vertical video, not a forced 16:9 crop.

When to reach for it

Use the reference mode when you want motion to follow a clip or the shot to sit on a beat. If you only need to keep one character's face consistent, other models on the canvas do that too — no need to detour.

The cover image above is a frame H3 generated.