MiniMax H3 Context Loop Trim
Cut the pinned frames off H3 output, picture and sound together
- images
- audio
- state
- images
- audio
- images_with_overlap
- overlap_frames
Every H3 chain scene comes out with a pinned overlap at the head - repeated context frames that have to come off before you concatenate scenes. Trim only the images and you've just shoved the soundtrack ahead of the picture by the same span: at 5 frames that's 208ms, inaudible on ambience but squarely offbeat on anything with a pulse. This node removes the leading pinned frames from both picture and sound so the clip is actually usable, and in 0.5 chains it resolves the per-scene blend from state instead of making you wire a constant.
What it does
Feed it the decoded images from the current H3 sample and a trim_frames count - normally the trim_frames output of MiniMax H3 Contex Loop Context, which knows what the encoder actually produced. Connect audio if the clip has sound; it gets trimmed by the matching duration. The fps (default 24) converts the frame trim into an audio duration and must match what you feed Create Video.
Outputs:
- images - the delivered frames with the repeated leading context removed.
- audio - trimmed by the same duration. With
match_tailon (the default), it also time-conforms small H3 audio-grid mismatches so duration equals frames/fps exactly - H3's rounded 40 Hz audio grid runs about 8ms off picture, and that error accumulates down a long chain. - images_with_overlap + overlap_frames - an optional blend-ready stream retaining only the final part of the repeated visual context, for external stitchers that need an overlap count.
In a 0.5 chain, the recommended wiring is to connect Current Shot's state output into state. Then the node reads the active scene's resolved blend directly from the Plan, so a per-scene value can never be overridden by a Plan default. When state is connected, the retain_overlap_frames integer is ignored (it's the legacy manual path for external stitchers).
Install and gotchas
cd ComfyUI/custom_nodes
git clone https://github.com/ethanfel/ComfyUI-MiniMaxH3-Contex-Loop.git
Restart, or install via ComfyUI Manager under "MiniMax H3 Contex Loop". No pip step; a current ComfyUI with native Add Guide (PR #15439) is expected, and ffmpeg on PATH is preferred. Models aren't bundled - mind the MiniMax H3 Community License territory restriction.
The classic mistake is trimming images only and letting the audio run long, which is exactly the desync this node prevents - so don't bypass it with a plain image slice. And if your audio drifts later in the chain, check match_tail; it's the fix for H3's ~8ms grid rounding, not a band-aid on your mux settings.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| images | IMAGE | Decoded images from the CURRENT H3 sample. The leading pinned overlap is removed. | |
| trim_frames | INT | 00–4096 | Connect trim_frames from MiniMax H3 Context Loop Context. In head mode this removes the repeated overlap; before mode supplies 0. |
| audioopt | AUDIO | Decoded audio for the same clip. Default sync_with_video removes the matching head duration; the explicit fresh-narration mode removes it from the tail instead. Leave unwired for silent clips. | |
| fpsopt | FLOAT | 24.0001–240 | Frame rate used to convert the trim into an audio duration. Must match what you feed Create Video. |
| match_tailopt | BOOLEAN | true | Time-conform small H3 audio-grid mismatches so duration equals frames/fps exactly without a silence tail. H3's rounded 40 Hz grid can differ from picture duration by about 8ms. |
| stateopt | H3_CHAIN_STATE | Recommended 0.5 chain route: connect Current Shot's state output. Loop Trim reads the active scene's resolved blend directly, so a Plan default can never override a per-scene value. | |
| audio_trim_modeopt | COMBO | sync_with_video | sync_with_video preserves A/V timing (default). fresh_narration_keep_start keeps the opening of off-screen narration and cuts the excess from the END to match delivered video length. This shifts sound relative to picture: NOT for lip-sync. Requires Current Shot state, fresh generated audio (no carry/source guide/lock), and match_tail. Leave room at the end for the removed duration; it can cut closing words. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| images | IMAGE | Delivered frames with the repeated leading context removed. |
| audio | AUDIO | Audio trimmed using audio_trim_mode and, when match_tail is enabled, fitted exactly to the delivered image duration. Synchronized mode privately carries decoded overlap to Segment Save for AV feathers; keep-start narration carries its timing choice without that overlap. Connect directly to Segment Save; no extra wire is needed. |
| images_with_overlap | IMAGE | Optional blend-ready image stream. When overlap_frames is positive, this retains only the final requested part of the repeated visual context before the delivered frames. Audio remains fully trimmed. |
| overlap_frames | INT | Number of repeated leading frames retained in images_with_overlap. Use as an overlap count in a compatible video stitcher. |