Nodes/ComfyUI-MinimaxH3DYTsc/minimaxH3DYTsc6.0 Conditioning
ComfyUI Node

minimaxH3DYTsc6.0 Conditioning

The node that picks i2v, fl2v or r2v for you

By 792877530-star·Created 6 days ago·Updated 6 days ago· 2
minimaxH3DYTsc6.0 Conditioning
  • clip
  • vae
  • audio_vae
  • first_frame
  • last_frame
  • reference_image_0
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_image_4
  • reference_image_5
  • reference_image_6
  • reference_image_7
  • reference_image_8
  • positive
  • latent
◄prompt►
◄width864►
◄height480►
◄length124►
◄ref_image_sizematch►

The problem it solves

ComfyUI got MiniMax H3 support on day zero - the team was handed the weights early, which is about as strong a signal as this ecosystem gives that a release is real. That put two official nodes in comfy_extras.nodes_minimax_h3: MiniMaxH3ImageToVideo (plain text, or with first/last keyframes) and MiniMaxH3ReferenceToVideo (reference images, source video, reference audio). Powerful, but you have to choose correctly, and the official nodes have quietly reshuffled their argument order across ComfyUI releases.

minimaxH3DYTsc6.0 Conditioning (class MinimaxH3DYTScConditioning) is a deliberately thin wrapper over both. You wire a CLIP, a VAE, a prompt and a size. If nothing reference-shaped is connected it calls MiniMaxH3ImageToVideo; the moment you connect a reference image, video or audio it switches to MiniMaxH3ReferenceToVideo. That's it. Reach for it when you're building a hand-wired H3 graph and don't want the pack's timeline UI - that's the Director node's job, and it's a much bigger commitment.

How it works

It imports the official classes at runtime and calls them with keyword arguments, so the signature reshuffles don't bite; it also checks once per class that the argument names it needs still exist and raises a readable error instead of letting a height // 16 blow up downstream. The latent it hands you is the same joint audio+video latent the official nodes build, not a video-only one.

The H3 framing matters here. This is MiniMax's 33B omni-modal model - text, image, video and audio as one input context, with stereo audio generated jointly rather than bolted on in a second pass. That's the actual reason to use H3 over Wan or LTX for a talking shot: the lip-sync and ambience are the same generation. Prompts go to a Qwen3-VL encoder in freeform, and reference material is addressed in the prompt by tag: <Picture 1>, <Video 1>, <Audio 1>.

Inputs worth touching

The required five are the whole story for a basic run. clip is the H3 text encoder loaded with type minimax; vae is the video VAE; prompt is freeform (no T5 prefix, no masterpiece incantations); width/height default to 864×480 in steps of 32; and length defaults to 124 frames - about five seconds at 24 fps.

length has a step of 17 because H3 snaps to a 17k+5 frame grid. 124 = 5s, 107 = 4.5s, 141 = 5.9s. Nothing in between survives.

The optional side is where the task type actually gets decided. first_frame and last_frame give you i2v / first-last-keyframe runs. reference_image_0 through reference_image_8 map to <Picture 1>…<Picture 9> - the author's tooltip is explicit that they keep native aspect and the H3 ref_image_size setting applies at encode time. ref_image_size is a two-way switch: match scales a reference to your canvas area, max scales it to the model's max short edge. audio_vae looks optional in the UI and isn't in practice - connect any reference material and the node refuses to run without it.

Wiring the outputs

Two outputs: positive goes to the sampler's positive input, latent goes to the sampler's latent. From there the official H3 chain does the work - MiniMaxH3SigmaShift for the sigma schedule (the official template uses 12 for video and 3 for audio), a KSampler, then decoding via LTXVSeparateAVLatent to split the joint latent into video and audio halves, VAEDecode for frames and the audio VAE path for the soundtrack. Assemble with CreateVideo at 24 fps.

Install

ComfyUI Manager → search minimaxH3DYTsc6.0, or:

cd ComfyUI/custom_nodes
git clone https://github.com/792877530-star/ComfyUI-MinimaxH3DYTsc

Restart ComfyUI. The pack's requirements.txt is light but real: opencv-python-headless, imageio-ffmpeg, scenedetect and comfy-kitchen>=0.2.34 (deliberately a floor, not a pin, so it doesn't fight ComfyUI's own pin).

What you actually need on disk: an H3 UNET in models/diffusion_models (minimax_h3_fl2va_* for the main model, minimax_h3_ref2va_* when you're using reference material), the Qwen3-VL H3 text encoder in models/text_encoders, the video VAE and minimax_h3_audio_vae_fp32.safetensors. The pack's built-in model card ships with int8 convrot names as defaults - those are fine if your ComfyUI is 0.27.0 or newer, which added native ConvRot INT8 (better quality than fp8 and faster on 20/30/40/50-series cards). On an older build, skip them.

Where people get burned

  • "requires ComfyUI official MiniMax H3 nodes" on load. Your ComfyUI predates the H3 extras. Update ComfyUI, restart, retry.
  • requires audio_vae the first time you connect a reference image. This is the single most common H3 trip-up: r2v/v2v/rv2v runs all need the audio VAE loaded, because audio is part of the conditioning, not an afterthought.
  • A reference video under 5 frames throws. That's roughly 0.2s at 24 fps; shorter clips get rejected outright.
  • GGUF weights. If your H3 files end in .gguf you need ComfyUI-GGUF installed and a restart - the pack looks the loader up in the live node registry and complains in plain language if it's missing.
  • Chinese error strings. A lot of this pack's error text is written in Chinese. Don't try to search the message; translate the gist and it's usually one of the four above.
CategoryMiniMaxH3

Inputs (19)

NameTypeDefaultDescription
clipCLIP—
vaeVAE—
promptSTRING—
widthINT86432–8192—
heightINT48032–8192—
lengthINT1245–3600—
audio_vaeoptVAERequired for r2v / v2v / rv2v / reference video+audio.
first_frameoptIMAGEOptional first keyframe (i2v / fl2v).
last_frameoptIMAGEOptional last keyframe (fl2v).
reference_image_0optIMAGEReference image for <Picture 1> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_1optIMAGEReference image for <Picture 2> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_2optIMAGEReference image for <Picture 3> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_3optIMAGEReference image for <Picture 4> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_4optIMAGEReference image for <Picture 5> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_5optIMAGEReference image for <Picture 6> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_6optIMAGEReference image for <Picture 7> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_7optIMAGEReference image for <Picture 8> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_8optIMAGEReference image for <Picture 9> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
ref_image_sizeoptCOMBOmatchReference image sizing for MiniMaxH3ReferenceToVideo.

Outputs (2)

NameTypeDescription
positiveCONDITIONING—
latentLATENT—