Nodes/ComfyUI/MiniMax H3 Reference to Video
ComfyUI Node Runs on cloud

MiniMax H3 Reference to Video

H3 with reference images, video and audio

By Comfy-Org·Created 4 years ago·Updated about 5 hours ago· 128,055
MiniMax H3 Reference to Video
    • VIDEO
    model
    seed42
    watermarkfalse

    Most video nodes take one input and run. This one takes a whole brief: up to nine reference images, three reference videos, three audio clips, and a prompt - and MiniMax H3 weaves them all into a single generation. Want the character from this image moving like the person in that video, over the music in this track? That's this node's whole personality.

    It's the most powerful of the H3 trio - and the most fiddly. It's a partner node, rendered on MiniMax's servers through Comfy's proxy and billed per second through your Comfy account, landed in ComfyUI core in late July 2026. The capability is real, but so are the constraints, and the node enforces them hard.

    How it works

    Everything is connected through the model dropdown, which holds the prompt plus three autogrow groups:

    • reference_images - up to 9, referred to in the prompt as "Image 1", "Image 2" and so on, in connection order.
    • reference_videos - up to 3, referred to as "Video 1".."Video 3"; each must be 2-15 seconds at 23.976-60 fps, 15 seconds total.
    • reference_audios - up to 3, "Audio 1".."Audio 3"; each 2-15 seconds, 15 seconds total.

    The prompt is where you bind them together: "Image 1 is the hero. Make them move like Video 1. Match the pace of Audio 1." That's the actual syntax - you reference the refs by their order number and the model understands.

    Also inside the dropdown: resolution (768P or 2K), ratio (adaptive or fixed), and duration (4-15s). Outside it: seed and the watermark toggle.

    The inputs that matter

    • model - prompt + all three reference groups + resolution/ratio/duration. This dropdown is the whole node.
    • seed - sent to MiniMax; similar, not identical, results on repeat.
    • watermark - AIGC watermark on/off, off by default.

    Output is a single VIDEO.

    Gotchas

    This node rejects sloppy input more than any sibling. Audio references can't be used without an image or video reference at all. Reference videos outside the 2-15 second window or fps range error out, and the 15-second total across all videos is a hard cap - plan the refs before you build. The per-second meter also includes extra charges past five reference images and for reference videos, so the price badge is worth actually reading. And the "Image 1 / Video 1" naming is load-bearing: if your prompt doesn't reference the inputs by order, you've paid for conditioning the model won't use.

    Categorypartner/video/MiniMax

    Inputs (3)

    NameTypeDefaultDescription
    modelCOMBOModel to use for video generation.
    seedINT420–4294967295Random seed. The same request with the same seed gives similar, but not guaranteed identical, results.
    watermarkBOOLEANfalseWhether to add an AIGC watermark to the video.

    Outputs (1)

    NameTypeDescription
    VIDEOVIDEO