Decode H3 Edit to One Image
Getting a Still Out of a Video Model
- samples
- vae
- image
- info
MiniMax H3 is a video model. You can ask it for a single still, and the naive way - one latent token, one frame - is genuinely bad, because a single frame has no temporal context around it. That's why this node exists: it decodes the whole short video stream the encoder set up and then picks the best frame for you.
If you've run the normal H3 chain, the usual ending is a VAEDecode and a SaveImage and you get a batch you have to eyeball. This replaces that with one image and a text line telling you which frame won.
The mechanism
Text Encode H3 Edit / Generate defaults to the recommended | 5-frame context -> 1 image profile, which asks H3 for a short clip rather than a still - the same short-context balance H3 Studio uses. The decoder then calls the H3 video VAE on the full latent, keeps the candidate frames the profile requested (5 by default, up to 20 on the maximum profile), and scores them.
The scoring is deliberately simple and it's the part worth understanding. It computes Laplacian variance as a sharpness proxy, grayscale standard deviation as contrast, and a pixel-clipping ratio as exposure, then weights them 0.70 / 0.20 / 0.10 and subtracts a penalty for how much each frame differs from its neighbours. So it's choosing the sharpest, best-exposed frame that also isn't mid-flicker. The info string reports the frame index and the winning score, which makes it obvious when the whole packet was mediocre rather than the selection being bad.
One special case: on the directed change | 39-frame settle -> 1 image profile, the encoder flags a settled-tail strategy, and the decoder only looks at the last five frames - zero-based 34–38, after the transformation is supposed to have finished and frozen. That's why re-pose, character-swap and camera-angle runs come back as a clean completed result rather than a half-morphed one.
Inputs and outputs
There are only two inputs, and both are required: samples (the sampled H3 edit latent from the encoder) and vae (the H3 video VAE). There's nothing to tune, which is either refreshing or frustrating depending on your mood.
Outputs are image - exactly one frame, not a batch - and info, the STRING with decoded frame count, the scored range and the winning index and score. Wire image into SaveImage or a preview; wire info into a text preview node if you want to see the numbers, or just leave it dangling.
Install
Manager → search the pack title, or:
cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMax-H3-Edit
Restart. Nothing extra to pip install - the pack ships with an empty dependency list on purpose. You need the H3 video VAE and a MiniMax H3 diffusion model loaded in the graph upstream.
The pack's mixed-reference still workflow (example_workflows/H3_Edit_Mixed_References.json) already ends in this decoder, so it's the fastest way to see it behave.
Where people get burned
- Feeding it a latent from somewhere else. The decoder expects an H3 video latent shaped
[B, 24, T, H, W]and it says so by name when it doesn't get one. It also reads frame-count metadata the encoder wrote onto the latent, so a latent built by another node loses that and you get default assumptions rather than a crash. - Decoding with a 4-D VAE output. Some VAE paths return
[T, H, W, C]instead of[B, T, H, W, C]; the decoder handles both, but it needs the encoded frame count to know whether it's looking at a real sequence or a single frame. Again - that metadata comes from this pack's encoder. - Expecting
experimental | true 1 frameto be fine. It isn't, and the profile name says so. It exists for interoperability with the older single-frame approach and for comparison, not for output. - Cranking
maximum. 20 candidates means 7 video latent tokens of sampling for one image.recommendedis the sane default; go up only when a difficult edit keeps landing on a soft frame.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| samples | LATENT | Sampled MiniMax H3 edit latent from Text Encode H3 Edit. | |
| vae | VAE | MiniMax H3 video VAE. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| info | STRING | — |