XinYu.
Navigation
Developers
MCPCLI
Language
← Back to News
New ModelAug 1, 2026

MiniMax H3 is live: reference with images, video, or audio

The 15-second clip above is straight text-to-video from H3 — an Apple-ad parody for a spring onion. No reference image, no titles added in post: the Chinese type you see was rendered by the model itself.

References aren't limited to images

Most video models only take images as reference. H3's reference mode accepts all three at once:

  • Images — lock a person, a product, a look
  • Video — hand it motion and camera movement to follow
  • Audio — hand it rhythm and mood

On the canvas, wire image, video and audio nodes into an H3 video node and switch to multi-reference — they all go in together. At least one image or video is required (audio alone won't do).

Three ways to run it

Text-to-video — wire nothing, just write the prompt.

Image-to-video — wire one image as the first frame; add a second as the last frame to move from A to B.

Multimodal reference — the one above: images, video and audio mixed.

Specs

Resolution2K (2560×1440) or 768P (1344×768); vertical is the transpose
Duration4–15 seconds, any whole second
Aspect21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16, plus "adaptive" in image and reference modes
SoundPicture and audio generated together, not dubbed afterwards
Price (2K)17.5 Xins/sec — 70 for 4s, 140 for 8s, 263 for 15s
Price (768P)12.5 Xins/sec — 50 for 4s, 100 for 8s, 188 for 15s

50% off through September 13, 2026. 768P at 6.25 Xins/sec (25 for 4s, 50 for 8s, 94 for 15s), 2K at 8.75 Xins/sec (35 for 4s, 70 for 8s, 132 for 15s). The panel shows the discounted price with the original struck through — no code needed. Regular pricing resumes after the promo.

The 768P tier came later and costs about 70% of 2K for the same length — worth using for test takes before committing to 2K.

"Adaptive" is its default: give it a vertical image and you get a vertical video, not a forced 16:9 crop.

When to reach for it

Use the reference mode when you want motion to follow a clip or the shot to sit on a beat. If you only need to keep one character's face consistent, other models on the canvas do that too — no need to detour.

The cover image above is a frame H3 generated.

Share