The 15-second clip above is straight text-to-video from H3 — an Apple-ad parody for a spring onion. No reference image, no titles added in post: the Chinese type you see was rendered by the model itself.
References aren't limited to images
Most video models only take images as reference. H3's reference mode accepts all three at once:
- Images — lock a person, a product, a look
- Video — hand it motion and camera movement to follow
- Audio — hand it rhythm and mood
On the canvas, wire image, video and audio nodes into an H3 video node and switch to multi-reference — they all go in together. At least one image or video is required (audio alone won't do).
Three ways to run it
Text-to-video — wire nothing, just write the prompt.
Image-to-video — wire one image as the first frame; add a second as the last frame to move from A to B.
Multimodal reference — the one above: images, video and audio mixed.
Specs
| Resolution | 2K (2560×1440) or 768P (1344×768); vertical is the transpose |
| Duration | 4–15 seconds, any whole second |
| Aspect | 21:9 / 16:9 / 4:3 / 1:1 / 3:4 / 9:16, plus "adaptive" in image and reference modes |
| Sound | Picture and audio generated together, not dubbed afterwards |
| Price (2K) | 17.5 Xins/sec — 70 for 4s, 140 for 8s, 263 for 15s |
| Price (768P) | 12.5 Xins/sec — 50 for 4s, 100 for 8s, 188 for 15s |
50% off through September 13, 2026. 768P at 6.25 Xins/sec (25 for 4s, 50 for 8s, 94 for 15s), 2K at 8.75 Xins/sec (35 for 4s, 70 for 8s, 132 for 15s). The panel shows the discounted price with the original struck through — no code needed. Regular pricing resumes after the promo.
The 768P tier came later and costs about 70% of 2K for the same length — worth using for test takes before committing to 2K.
"Adaptive" is its default: give it a vertical image and you get a vertical video, not a forced 16:9 crop.
When to reach for it
Use the reference mode when you want motion to follow a clip or the shot to sit on a beat. If you only need to keep one character's face consistent, other models on the canvas do that too — no need to detour.
The cover image above is a frame H3 generated.