ComfyUI Node Runs on cloud

H3 Control

Drive every segment of an extend loop from one control video

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
H3 Control
  • model
  • model_patch
  • vae
  • window
  • control_video
  • mask
  • source_video
  • model
  • report
◄strength1.00►
◄start_percent0.00►
◄end_percent1.00►

The awkward part of any scene-by-scene video workflow is that each pass knows about its own segment and nothing else. So structural control - pose, depth, canny, lines - usually means building a separate control video per segment and remembering which frames belong to which. This node refuses to do that. You wire the whole video's control frames once, and each pass reads the frames that land on its own window.

How it works

H3 Control goes between H3 Extend Window's model output and the guider, and it works off the window's place on the finished clip rather than off any counter you maintain. If the window knows it starts at frame 340 of the finished video, the node encodes the control frames around that point to match, and the segment follows the control video in step. A latent with no place - an empty H3 latent for a single shot, say - reads from frame 0.

The control frames are encoded with the video VAE and stacked into a control latent alongside a visibility channel and, when you're inpainting, masked source channels: 24 control channels, one visibility channel, 24 source channels. The model is then wrapped by the patch before it reaches the guider, so what you're wiring downstream is an already-conditioned model.

The weights are yours to find. The pack does not ship the H3 Fun ControlNet-Union patch - you need a minimax_h3_fun_controlnet_union file in models/model_patches and core Load Model Patch to load it.

The inputs that matter

  • model - this segment's H3 model, from H3 Extend Window. Wire it here, then this node's model output goes to the guider, not the original.
  • model_patch - the ControlNet-Union patch, as above.
  • vae - the H3 video VAE. It encodes the control frames.
  • window - from H3 Extend Window. Its place on the finished clip is what picks the control frames.
  • strength - the usual control-net dial. 0.0 passes the model straight through, 1.0 is as trained, 0.5 is a looser follow your prompt can push against, and 1.5 is stricter than the model was trained for. Start at 1.0 and move it in 0.1s.
  • start_percent / end_percent - where in the noise schedule control applies. end_percent is the useful one: pulling it to 0.6 releases control for the final steps so fine detail comes from your prompt instead of a depth map.
  • Optional control_video - the whole video's control frames at 24 fps, frame 0 lining up with the finished clip's frame 0. It's cropped to the window's canvas and a short video holds its last frame rather than erroring, which is forgiving in a nice way.
  • Optional mask and source_video - 1 in the mask means regenerate, 0 means keep what's in source_video. The mask is the length of the whole video at 24 fps, one mask frame holds for every video frame, and it does nothing without a source video.

Two outputs: the patched model for the guider, and a report that says which control frames this window read, at what strength, and whether it's inpainting.

Installing it

ComfyUI Manager, search WAS Node Suite v3 - or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+. The pack installs nothing and fetches nothing; it names the file it wants instead. Its 2023-era reputation for dragging in OpenCV and breaking on ComfyUI updates is a v2 story - this is the third generation of the pack.

You'll need the H3 model set in place, plus the control patch. And a fair warning: MiniMax H3's local weights ship under a Community License whose grant excludes the US, EU, UK and South Korea. Nothing technical stops you; the licence is the constraint.

Where it goes wrong

Control from the wrong part of the video. Frame 0 of your control video is frame 0 of the finished clip, not of the current segment. If your control frames are a fifteen-second clip and your video is sixty seconds, scenes four onward get the held last frame.

The mask does nothing. It needs source_video, and it's a whole-video mask. Handing it a single-frame mask regenerates one frame and then behaves as though the rest is unmasked.

Everything looks like the control video and nothing like your prompt. Drop strength, or pull end_percent down so the last steps are the prompt's. This is the usual control-net trade, made sharper here because the video model already has a strong opinion about what happens next.

Patch loaded, no effect. Check that this node's model output is what feeds the guider. The patch is applied here; bypassing the node leaves you a perfectly normal, perfectly uncontrolled H3 graph.

CategoryWAS Suite/Latent/Video

Inputs (10)

NameTypeDefaultDescription
modelMODELThis segment's MiniMax H3 model, from H3 Extend Window's model output.
model_patchMODEL_PATCHThe H3 Fun ControlNet-Union patch, from Load Model Patch set to a `minimax_h3_fun_controlnet_union` file in models/model_patches.
vaeVAEThe H3 video VAE, which encodes this segment's control frames.
windowLATENTThe window this segment samples, from H3 Extend Window. Its place on the finished clip picks the control frames; a latent with no place reads from frame 0.
strengthFLOAT1.000–10How hard the control frames steer: `0.0` = off, model passed through; `1.0` = as trained; `0.5` = a looser follow the prompt can bend; `1.5` = stricter.
start_percentFLOAT0.000–1Where in the noise schedule control starts: `0.0` = from the first step; `0.2` = a little later, leaving the opening steps to the prompt.
end_percentFLOAT1.000–1Where in the noise schedule control stops: `1.0` = to the last step; `0.6` = released for the final steps, so fine detail follows the prompt rather than the control frames.
control_videooptIMAGEThe whole video's control frames at 24 fps, frame 0 lining up with the finished clip's frame 0: pose, depth, canny, lines or any map the patch takes. Cropped to the window's canvas; a short video holds its last frame.
maskoptMASKWhere to regenerate, the whole video's length at 24 fps: `1` = regenerate, `0` = keep source_video. One frame holds for every frame. Needs source_video.
source_videooptIMAGEThe video the mask was drawn on, the whole length at 24 fps. Kept outside the mask and regenerated inside it. Read only with a mask.

Outputs (2)

NameTypeDescription
modelMODELThe segment's model steered by its control frames, for the guider.
reportSTRINGWhich control frames the window reads, at what strength, and whether it inpaints.