Add Masked Guide for MiniMax H3
Pin a face, a clip or a soundtrack to any frame of an H3 video — and make it stick
- positive
- latent
- vae
- audio_vae
- image
- audio
- mask
- positive
- LATENT
What it is, and why it exists
MiniMax H3 is the 33B omni-modal video model MiniMax opened in August 2026: text, images, video and audio go into one context, a 4–15 second clip at up to 2K with synced stereo sound comes out. ComfyUI got day-zero support, and core ships MiniMaxH3AddGuide, which pins an image or clip at a given frame.
This node is a fork of it that adds the thing people kept asking for - a mask - plus hard enforcement. That difference matters because a pinned guide in the conditioning alone steers rather than holds. You give the model a condition row saying "this is what frame 40 looks like", and it does something inspired by that - often with flicker on the anchored region. So this fork also rewrites the latent: anchored patches go into the target latent and get marked preserved via a per-element noise_mask, so the sampler re-injects the guide content every step. That's the latent-inpaint path. Masked regions hold; everything else generates freely.
That's what you want for mid-clip keyframes, holding a face through a shot, or dropping a real soundtrack onto one segment of a clip.
How the anchoring works
H3's latent doesn't divide into even frames: its tokens cover 1, 4, 8, 8, 8 frames on repeat, so valid clip lengths are 17k + 5. The node maps your guide onto that grid by maximum overlap.
With an image connected, the frames are resized to the latent's resolution, encoded by the video vae, and appended to the positive conditioning's internal minimax_keyframes list, which drives the model's per-row condition timesteps.
The mask matters. It's resized to follow your guide frames, then max-pooled onto the DiT's 32-pixel patch grid - a patch anchors if any pixel inside it is masked, so your mask edge rounds up to the nearest patch, never shrunk or shifted. With mask_threshold at its default 0.5 the mask is binarised after that resize, which is what kills the feathered grey halo at a guide's border. Push the threshold below 0 and grey values become partial anchor strength instead - the model half-follows the guide there.
Audio, if you connect it, is resampled to the audio VAE's rate and clipped to whatever video remains after your frame index.
Inputs and outputs that matter
Required: positive, latent, frame_idx, mask_threshold.
frame_idx is the frame the image - or the first frame of your clip - is anchored to. Negative values count from the end, so -1 parks an image on the last frame; no separate "last frame" node needed.
Optional: vae (needed whenever an image is connected), audio_vae (whenever audio is connected), image, audio, mask. The mask applies to the image guide only - wiring a mask with no image is an error.
Outputs are positive and a LATENT: conditioning to the sampler's positive, and the latent to the sampler's latent slot instead of the empty H3 latent. The hard protection lives in that second output - skip it and you're back to a flickering guide.
Installing it
Manager: search ComfyUI-SA-Nodes-QQ. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/siraxe/ComfyUI-WanVideoWrapper_QQ.git
Restart ComfyUI, and delete any leftover wanwrapper_qq folder from the old repo name. requirements.txt is empty, so nothing extra installs - but you need a ComfyUI recent enough to have H3 in core: the node imports comfy.ldm.minimax.model, nested tensors and comfy_api.latest, so an older install fails at import and takes the whole pack with it. torchaudio is imported lazily on the audio path only.
Where people get burned
Your clip gets silently cropped. Multi-frame images are cut to the 17k + 5 grid: 30 frames becomes 22, 60 becomes 56. Under 5 frames only the first image is used, so a 3-frame batch behaves like a still.
"does not fit in the video" errors are that frame grid again. A 5-frame guide at frame_idx=0 in a 22-frame latent is fine; the same guide at frame 18 isn't. Negative indices are the safe way to anchor at the end.
Wrong latent shape. Feed it anything but H3's nested video+audio latent and it stops with "expects a MiniMax H3 AV latent". This is a fork for one model, not a general guide node.
Hardware and licence. The open weights are ~42.5 GB at full precision with no published consumer floor - don't plan a 12GB card around it until you see it run. And the H3 Community License excludes the US, EU, UK and South Korea from its Applicable Territory, outputs included - the base model's problem rather than this node's, but yours either way.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| positive | CONDITIONING | — | |
| latent | LATENT | MiniMax H3 AV latent. A single frame latent (Fizgig H3 Still Latent) is detected and handled as a still: one frame, so only frame 0 exists and only the guide's first frame is used. | |
| frame_idx | INT | 0-9999–9999 | Frame index to anchor the image or the clip first frame at. Negative values are counted from the end of the video. A single-frame latent only holds frame 0, so anything else is rejected. |
| mask_threshold | FLOAT | 0.50-1–1 | Binarize the mask above this value for a hard anchor edge (recommended: kills feathering at the mask border). Set below 0 to keep the mask soft, with gray values giving partial guide strength. |
| vaeopt | VAE | Video VAE, needed when an image is connected. | |
| audio_vaeopt | VAE | Audio VAE, needed when an audio is connected. | |
| imageopt | IMAGE | Image or video frames to anchor. Multi-frame batches are anchored as a clip and cropped down to the model valid clip lengths: 5, 22, 39... (17k + 5) frames. Batches shorter than 5 frames use only the first image. With a single-frame latent only the first frame is used. | |
| audioopt | AUDIO | Soundtrack to anchor starting at the same frame index, cropped to the video remaining duration. | |
| maskopt | MASK | Mask over the image guide: 1 keeps the guide anchored, 0 lets the model generate those regions freely. A single-frame mask applies to every frame of a guide clip; a multi-frame mask maps onto the clip by relative time. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | — |
| LATENT | LATENT | — |