MiniMax H3 Plan Transitions
Let a language model pick your cuts
- vlm_clip
- transitions
- report
The difference between an H3 video that feels like a film and one that feels like a slideshow is almost entirely in the transitions. Whether scene 3 carries on the shot, cuts hard, or cuts while keeping the sound changes how the thing reads - and picking eight of those by hand, twice a night, is exactly the kind of job people stop doing properly. This node reads your scenes with a language model and picks them for you.
How it works
You hand it the scene prompts in order, plus a vlm_clip - a language model loaded through ComfyUI's own CLIP path - and it answers with one decision per scene from the second onward: a cut, a carry of the same shot, a cut that keeps the sound, or a cut that keeps the cast, with or without the sound. The choices map onto the same transition vocabulary MiniMax H3 Conditioning's rows use (carry, handoff, reference (video), reference (sample), cut, and the audio-carrying variants), and each entry also says why, which is genuinely useful when one of them is wrong.
It's a text node. Nothing here samples video and no model is loaded by a render - this decides, the loop renders. The seed picks which draw of the model's choices you get; keep it and you get the same plan back, which is how you compare two prompts without also comparing two sets of transitions.
If you're using the Prompt Timeline window, you can skip the wiring entirely: Plan transitions in the timeline bar runs this node and writes each row's transition and overlap for you.
The inputs that matter
vlm_clip- from Load CLIP.qwen3vl_8b_fp8_scaled.safetensorsorqwen_3_8b_fp8mixed.safetensorsare the names the node suggests; put them inmodels/text_encoders.scenes- the prompts in order, divided by a line of dashes exactly as MiniMax H3 Scene Writer'spromptsoutput produces them. You can also paste the Scene Writer's scenes JSON.seed- which draw. Same seed and inputs, same plan.- The generation dials:
temperature(0.3 by default - low, because you want the likeliest reading rather than a creative one),top_k,top_p,min_p,repetition_penalty,thinkingfor reasoning models such as Qwen3, andfast_decode.
Two outputs. transitions is JSON, one entry per scene from the second, with the scene number, the transition, the continuity and overlap it implies, and the reason. report is a count of how many of each were chosen - a quick sanity check that you didn't get fourteen hard cuts.
Installing it
It ships in WAS Node Suite v3: ComfyUI Manager, search WAS Node Suite v3, or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+, and the pack installs nothing - no llama.cpp, no server, no separate Ollama process. The model is a ComfyUI CLIP, loaded by core Load CLIP.
That does mean real VRAM is in play. A Qwen3-VL 8B in fp8 is a few gigabytes resident, which is the standard local-LLM-in-ComfyUI arithmetic: you're budgeting for a second model beside the video model. If you're tight, a smaller quant is the usual answer, and the pack's fast_decode option keeps the decode on the fixed-cache path where the model allows it.
Where it goes wrong
Garbage or refusals in the output. Turn temperature down before anything else. 0.3 is the default for a reason; the node wants a deterministic picker, not a novelist.
Every transition is cut. Either the model can't see continuity to preserve, or the scenes share none. If scene 4 is meant to continue scene 3's shot, its prompt has to say so - the planner reads prompts, not your intentions. References and keyframes declared in your asset chain are part of what it has to work with, so wire the MiniMax H3 Asset chain in before planning.
"Nothing happened" after execution. Correct - it produced text. Look at transitions (use a preview or a show-text node) or at the rows in the Prompt Timeline, not at the canvas.
thinking: true and it takes forever. Reasoning models think before answering, and that's tokens on your card. It's a good switch for a two-scene plan and a bad one for a twenty-four-scene plan.
The format matters. Dashes between prompts, or the Scene Writer's JSON - not a numbered list you typed by hand. Feed it the wrong shape and it plans transitions for whatever it thinks the scenes are.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| vlm_clip | CLIP | The language model that reads the scenes, from Load CLIP, as `qwen3vl_8b_fp8_scaled.safetensors` or `qwen_3_8b_fp8mixed.safetensors`. | |
| scenes | STRING | The scene prompts in order, divided by a line of dashes as MiniMax H3 Scene Writer's prompts output, or its scenes JSON. | |
| seed | INT | 00–18446744073709550000 | Which draw of the model's choices to take, as `0` or `42`. |
| temperatureopt | FLOAT | 0.300.01–2 | How adventurous each pick is. `0.3` keeps to the likeliest reading, `0.8` varies it. |
| top_kopt | INT | 640–1000 | Pick only among this many likeliest tokens. `0` turns the limit off. |
| top_popt | FLOAT | 0.950–1 | Pick only among the likeliest tokens whose chances add up to this. `1.0` turns it off. |
| min_popt | FLOAT | 0.050–1 | Drop any token less likely than this fraction of the likeliest one. `0` turns it off. |
| repetition_penaltyopt | FLOAT | 1.050–5 | Above `1.0` makes a token already used less likely again. `1.0` turns it off. |
| thinkingopt | BOOLEAN | false | `true` lets a model that reasons, such as Qwen3, think before answering. |
| fast_decodeopt | BOOLEAN | true | `true` decodes on the fixed cache graph path where the model allows it; `false` runs as core Generate Text. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| transitions | STRING | One entry per scene from the second, as JSON: scene number, transition, the continuity and overlap it sets, and why. |
| report | STRING | How many of each transition were chosen. |