Nodes/comfyui-minimax-h3-audio-T8/H3 Audio Refine · Match Frozen Video Frames (T8 EXP)
ComfyUI Node

H3 Audio Refine · Match Frozen Video Frames (T8 EXP)

Catching A Wrong-Length H3 Checkpoint Before It Costs You A Render

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Audio Refine · Match Frozen Video Frames (T8 EXP)
  • av_latent
  • av_latent
  • video_width
  • video_height
  • video_frame_count
◄expected_video_frame_count124►

Resuming an audio refine against a frozen H3 checkpoint is a great way to iterate on sound without regenerating video. It's also a great way to waste twenty minutes on a tail pass attached to the wrong clip. This node is the cheap insurance: it fails before the refine if the frozen AV doesn't have the frame count your graph expects.

Two inputs. Four outputs. No GPU work. It's the least glamorous node in the pack and one of the more useful ones.

How it decides

H3's video latent doesn't map to frames one-to-one. Its geometry runs on a 17-frame grid, and this node reconstructs the real frame count from the latent's temporal dimension: a latent of 1 is 1 frame, otherwise the count is 5 + 17 × ((latent − 2) / 5). A latent that doesn't land on that grid is rejected outright as non-native H3 geometry.

That grid is why the pack talks about 5 / 22 / 39-frame continuation contexts so much, and why 124 frames - the default for expected_video_frame_count, about 5.17 seconds at 24fps - is such a common number. It's 5 + 17 × 7.

Valid counts, if you're planning a run: 5, 22, 39, 56, 73, 90, 107, 124, 141 and so on. Pick anything else and you get an error rather than a render that's one frame off in a way you notice after the fact.

The guard clamps its expectation between 1 and 3600 frames, so it won't let you declare a four-minute clip and walk away.

The outputs

av_latent passes your latent straight through, so you can drop this node inline between the loader and the refine without rewiring anything. video_width and video_height report the frame geometry the latent implies - the node derives them from the latent's spatial dimensions rather than trusting any widget, which is exactly what you want when the whole point is "is this the file I think it is". video_frame_count is the reconciled count.

Those three INT outputs are handy for feeding downstream size fields, and even handier as a sanity print on a resume graph.

Where it goes

Inline, immediately after MiniMaxH3AudioRefineFrozenFirstPassLoadEXPT8 and before the stage bind. The loader already validates the checkpoint's manifest, file SHA-256 and audio_refine_firstpass ID; the guard adds the one thing a hash can't tell you - whether this graph's expected duration matches the file, which is exactly the mistake you make when you've got three checkpoints in output/MiniMaxH3/latent_checkpoints and the file names stopped being memorable an hour ago.

By design, empty evidence fails before any tail sampling. There's no "close enough" branch, no auto-trim. The pack's stance is that guessing here is how you end up with a 73-frame tail welded onto a 124-frame video and no clear error to explain it.

Installing

Method one: ComfyUI Manager, search MiniMax H3 Audio T8. Method two, and the one I'd use if you care about matching a specific version:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Then fully quit ComfyUI and relaunch - not a browser refresh, an actual restart - because node classes register at startup. There's no dependency step: the pack's requirements file deliberately declares nothing beyond what ComfyUI already ships, so installing it can't clobber your Torch or CUDA build. You will need a recent ComfyUI with native H3 support, and the weights are yours to supply: H3 diffusion model into models/diffusion_models, Qwen text encoder into models/text_encoders, video and audio VAEs into models/vae.

Bugs and traps

  • Setting expected_video_frame_count to a number off the 17-frame grid. You'll get a geometry rejection even if the checkpoint is fine, because the expected count can never be matched. Use the grid.
  • Treating a pass here as proof the checkpoint is good. It proves the length matches. Quality, audio, provenance - all elsewhere.
  • Putting it after the sampler. It's a pre-flight check. Downstream of the refine it's just decorative.
  • Assuming it auto-detects. It doesn't; you declare the count and the node enforces it. That's the point.
CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (2)

NameTypeDefaultDescription
av_latentLATENT—
expected_video_frame_countINT1241–3600—

Outputs (4)

NameTypeDescription
av_latentLATENT—
video_widthINT—
video_heightINT—
video_frame_countINT—