MiniMax H3 joins the video models: text-to-video, image-to-video (with optional last frame), and multimodal reference (image + video + audio). 2K, 5–15s, native audio, 17.5 Xins/sec. See the announcement.