Nodes/MiniMax H3 Planner/H3 Track Slice
ComfyUI Node

H3 Track Slice

Give the sampler the exact window of the track this segment needs

By AIJigyasa·Created 23 days ago·Updated 20 days ago· 4
H3 Track Slice
  • audio
  • shot
  • audio
  • start
  • end
  • info
◄pad_shorttrue►
◄offset_seconds0.00►

H3's ref_audio_0 input takes one piece of audio per generation. That's fine for a single clip and useless for a four-minute music video: segment seven needs the window from 43.2s to 49.8s, and nothing in ComfyUI computes that for you. H3 Track Slice does - it reads the segment's own audio_start off the shot token and hands the sampler exactly the slice it should hear.

It's a two-wire node and it's the difference between a video that lip-syncs and one that drifts.

One subtle decision worth knowing

It slices render_duration, not target_duration. That matters: H3 renders slightly longer than planned because of the frame ladder, and if the audio stopped at the planned length, the last half-second of every clip would be picture with silence under it. By slicing the full rendered length, the stitcher's trim removes picture and sound overhang together, and the join stays clean.

That's the same reasoning behind the plan-vs-render duration split everywhere else in the pack: plan in musical or authored time, render on the ladder, trim on the way out.

Inputs and outputs

Required:

  • audio - the full track. Straight off the H3 Cast Board's audio output is the intended path, and the node's own tooltip says so.
  • shot - the token from H3 Shot Dispatcher. Without it there's no segment, no audio_start, and no window to compute. This is the wire that goes around the sampler.
  • pad_short - pad with silence when the track runs out. Leave it on unless you'd rather find out, loudly, that your last segment runs past the end of the song.

Optional offset_seconds nudges the window, and it takes negative values too - handy when a shot wants to land half a beat before its cut, or when your DAW and your file metadata disagree about where zero is.

Outputs: audio goes to the sampler's ref_audio_0. start and end are the actual seconds used, which is how you sanity-check the window without doing arithmetic. info prints the segment id, the window, how much of the total track it covers, and a note when it had to pad or came up short.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/AIJigyasa/ComfyUI-H3-Planner

Restart, or install MiniMax H3 Planner from ComfyUI Manager. No extra Python packages for this node - it slices tensors. The pack wants ffmpeg on PATH (or pip install imageio-ffmpeg) because stored clips get encoded, and a ComfyUI with the MiniMax H3 nodes, since the pack drives MiniMaxH3ReferenceToVideo rather than sampling.

Where people get burned

Feed it the whole track, not a trimmed clip. This is the one that catches people coming from a single-clip graph. The pack's own workflow notes flag it: if your audio loader still has start_time 73.03 / end_time 83.12 baked in from when the graph made a 10-second clip, then audio_start becomes relative to that 10-second slice and every segment gets the wrong window. Set the loader to hand over the full track and let the per-segment audio_start pick its own window - that's the entire point of this node.

Silence at the end of a segment - either the track ran out (pad_short is doing its job, and info says so), or your audio_start plus duration overshoots the file. Check start/end against the track length.

Audio that doesn't line up with the cut is usually a stitcher problem, not this one. If H3 Stitch Timeline has trim_to_plan off, the plan/render mismatch accumulates and the track slides against the picture. Leave the trim on, and use audio: from track on the stitcher for a music video - it lays the original track over the whole video at full quality rather than joining twelve separately-generated audio snippets.

CategoryH3 Planner

Inputs (4)

NameTypeDefaultDescription
audioAUDIO—
shotH3_SHOT—
pad_shortBOOLEANtruepad with silence when the track runs out
offset_secondsoptFLOAT0.00-600–600—

Outputs (4)

NameTypeDescription
audioAUDIO—
startFLOAT—
endFLOAT—
infoSTRING—