Nodes/ComfyUI-MinimaxH3DYTsc/minimaxH3DYTsc6.0 PlannerConditioning
ComfyUI Node

minimaxH3DYTsc6.0 PlannerConditioning

The same H3 conditioning node, plus one useful string

By 792877530-star·Created 6 days ago·Updated 6 days ago· 2
minimaxH3DYTsc6.0 PlannerConditioning
  • clip
  • vae
  • audio_vae
  • first_frame
  • last_frame
  • reference_image_0
  • reference_image_1
  • reference_image_2
  • reference_image_3
  • reference_image_4
  • reference_image_5
  • reference_image_6
  • reference_image_7
  • reference_image_8
  • positive
  • latent
  • task_mode
◄prompt►
◄width864►
◄height480►
◄length124►
◄ref_image_sizematch►

What it is

minimaxH3DYTsc6.0 PlannerConditioning (class MinimaxH3DYTScPlannerConditioning) is the conditioning node from this pack with one extra output. Same inputs, same wiring, same behaviour - the class even inherits its input definition from the plain Conditioning node, so you can swap one for the other in a graph without re-plumbing anything.

The addition is a third output, task_mode, a plain STRING.

If you're doing a hand-wired MiniMax H3 graph, the interesting part is unchanged from the base node: it auto-routes to ComfyUI's official MiniMaxH3ImageToVideo when nothing reference-shaped is connected, and to MiniMaxH3ReferenceToVideo when a reference image, video or audio is. positive goes to your sampler's positive input, latent to its latent input - a joint audio+video latent, matching H3's whole premise of generating stereo sound with the picture instead of bolting it on afterwards.

What the string actually says

Nothing mysterious. The node counts what you connected, infers the task, and formats a human sentence:

r2v - Reference-to AV - subject images (+ optional tags) in prompt (MiniMax H3) (~3 ref image(s), 0 ref video(s))

With no references connected you get t2v - Text-to AV (no keyframes or references) (MiniMax H3) instead. There are six task descriptions in the pack's vocabulary (t2v, i2v, fl2v, r2v, v2v, rv2v), and the inference itself is blunt: any reference connected means r2v, nothing connected means t2v. It is a label describing what you wired, not a plan and not a prediction. It can't tell you an i2v run from a plain t2v run.

So why does it exist? Because a string output is the cheapest way to get that state somewhere useful. Route it into a Show Text or note node to sanity-check which path your graph will take before you burn ten minutes on a 124-frame render. Log it. Use it as a title for a batch. Or feed it to your own node that switches behaviour on the mode - the name PlannerConditioning is the author signalling that this is the one to grab if a downstream node of yours wants to know what kind of H3 run it's looking at.

Which one to pick

Take this one if you have a use for a text string. Otherwise use MinimaxH3DYTScConditioning: it's the same node with one fewer dangling output, and every connection you don't make is one fewer thing to mis-wire. Both are in category MiniMaxH3 and both live in the same pack.

The inputs you'll set are the same five required ones - clip (H3 text encoder, CLIPLoader type minimax), vae (video VAE), prompt (freeform Qwen3-VL text, with <Picture N> / <Video K> / <Audio J> tags when you use references), width/height (864×480 default, steps of 32) and length (124 frames ≈ 5s at 24fps, on H3's 17k+5 grid). Then audio_vae, which is required the moment you connect any reference material even though the UI calls it optional, and optionally first_frame / last_frame for keyframe runs or reference_image_0…reference_image_8 for <Picture 1>…<Picture 9>.

Install

Identical to the rest of the pack - it's not a separate download:

cd ComfyUI/custom_nodes
git clone https://github.com/792877530-star/ComfyUI-MinimaxH3DYTsc

Restart, then look under MiniMaxH3 in the node search. Manager users can search minimaxH3DYTsc6.0. You need a ComfyUI new enough to ship comfy_extras.nodes_minimax_h3, plus the H3 UNET, the Qwen3-VL H3 text encoder, the video VAE and minimax_h3_audio_vae_fp32.safetensors on disk.

Gotchas

Same as its sibling, and they're worth repeating because they account for nearly every "it won't run" post about H3 in ComfyUI. Missing official nodes means your ComfyUI needs updating. A reference run without audio_vae is refused by design. Reference videos need at least five frames. GGUF weights need ComfyUI-GGUF installed and ComfyUI restarted before the loader is visible. And most of this pack's error text is in Chinese - the messages are specific and accurate, but pasting them into a search box won't help you.

CategoryMiniMaxH3

Inputs (19)

NameTypeDefaultDescription
clipCLIP—
vaeVAE—
promptSTRING—
widthINT86432–8192—
heightINT48032–8192—
lengthINT1245–3600—
audio_vaeoptVAERequired for r2v / v2v / rv2v / reference video+audio.
first_frameoptIMAGEOptional first keyframe (i2v / fl2v).
last_frameoptIMAGEOptional last keyframe (fl2v).
reference_image_0optIMAGEReference image for <Picture 1> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_1optIMAGEReference image for <Picture 2> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_2optIMAGEReference image for <Picture 3> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_3optIMAGEReference image for <Picture 4> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_4optIMAGEReference image for <Picture 5> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_5optIMAGEReference image for <Picture 6> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_6optIMAGEReference image for <Picture 7> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_7optIMAGEReference image for <Picture 8> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
reference_image_8optIMAGEReference image for <Picture 9> in prompt (r2v). Native aspect; H3 ref_image_size applies at encode time.
ref_image_sizeoptCOMBOmatchReference image sizing for MiniMaxH3ReferenceToVideo.

Outputs (3)

NameTypeDescription
positiveCONDITIONING—
latentLATENT—
task_modeSTRING—