ComfyUI Node Runs on cloud

H3 Soundtrack

One music bed under a video that was stitched together from scenes

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
H3 Soundtrack
  • window
  • audio
  • audio_vae
  • window
  • report
◄offset_seconds0.000►
◄strength1.00►
◄release8►
◄past_endmodel sound►

MiniMax H3 generates sound with its picture, which is lovely until you want a score. A video built scene by scene gets new audio every scene, and every cut is a fresh chance for the music to reset. This node fixes that: you hand it one long track - a music bed, a recorded voice-over, a dialogue pass - and each segment samples its picture under its own slice of that same track. Drop it between H3 Extend Window's window output and the sampler's latent input.

How it works

The window carries a stretch of the video, and part of what it carries is a place on the finished clip. The node uses that place to work out which slice of your track belongs under this window, encodes that slice into the window's audio latent, and hands the window straight back out. Sound the window already carries from earlier segments is left alone; offset_seconds shifts the whole track against the video so you can start the music two and a half seconds in, or skip its first ten seconds.

The encoding is the expensive part, and it happens once. The track is encoded at the H3 audio rate of 40 latent steps per second, in pieces with the seams dropped, with a second of silence appended after the end, and the result is cached - every later segment reuses it. A three-minute track is not re-encoded ten times because you have ten scenes.

The inputs that matter

  • window - from H3 Extend Window, or an empty H3 AV latent if you're scoring a single-shot video. It comes back out as window, so the sampler wiring is unchanged.
  • audio - the whole track, as Load Audio hands it over. Any sample rate; a mono track plays on both sides. Same track into every segment; that's the point.
  • audio_vae - the H3 audio VAE, which you load with core Load VAE set to minimax_h3_audio_vae_fp32. Not the video VAE.
  • offset_seconds - where the track starts relative to the finished video. 0 lands it on frame 1, 2.5 lets the model's own sound open the video before the music comes in, -10 trims the track's first ten seconds.
  • strength - how firmly the track is held, 0 to 1. At 0 the window passes through untouched, which is also how you A/B whether the node is helping. At 0.5 the sound is guided rather than locked and can drift from the track; at 1.0 you get the track.
  • release - audio steps of ease at the track's start, end and any loop point, at 40 a second. 8 is the default 0.2 seconds. 0 gives you a hard edge, which is audible exactly where you'd expect.
  • past_end - what plays once the track runs out before the video does: model sound, silence, or loop with the seam eased over release.

Two outputs, window and a report that tells you which stretch of the track landed under which frames. The report is the underrated part; when the music feels nearly right, it's usually because your offset is a beat off, not because strength needs fiddling.

Installing it

It ships in WAS Node Suite v3: ComfyUI Manager, search WAS Node Suite v3, or:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

ComfyUI 0.14.0+ and Python 3.10+, and the pack installs nothing. What you will need is the whole H3 model set: a transformer in models/diffusion_models, the video and audio VAEs in models/vae, and the text encoder in models/text_encoders. H3 weights are under MiniMax's Community License, whose grant excludes the US, EU, UK and South Korea - worth a look before you build a pipeline on it.

Where it goes wrong

The music starts late, or early, by a consistent amount. That's offset_seconds and your own ears disagreeing about where the video begins. Use the report; it prints the frame ranges.

Strength above 0.5 and the track sounds like mud. You're asking H3 to generate audio that matches a foreign waveform it isn't trained to copy exactly. Keep it at or near 1.0 for the sound to actually be your track, and drop toward 0 if you only want it as a hint.

Nothing changed at all. Check strength - at 0 the node is a pass-through by design, and a workflow saved with it there looks like the node did nothing.

This isn't a foley node. If you wanted sound generated for a silent clip, you want MMAudio or a similar video-to-audio pass, not this. H3 already made its own sound; this node is about replacing a stretch of it with something you brought.

CategoryWAS Suite/Latent/Video

Inputs (7)

NameTypeDefaultDescription
windowLATENTThe window to sample, from H3 Extend Window's window output, or an Empty MiniMax H3 AV Latent for a single shot. Passed on to the sampler's latent input.
audioAUDIOThe whole track, as Load Audio's output: a music bed or a recorded dialogue. Any sample rate; a mono track plays on both sides. The same track goes into every segment.
audio_vaeVAEThe H3 audio VAE, as Load VAE set to `minimax_h3_audio_vae_fp32`. It encodes the track once and every later segment reuses that.
offset_secondsFLOAT0.000-36000–36000Where the track starts on the finished video, in seconds: `0` = with the first frame; `2.5` = 2.5 s in, with the model's own sound before it; `-10` = skips the track's first 10 s.
strengthFLOAT1.000–1How firmly the track is held under the new picture: `0` = ignored, the window passes through; `0.5` = guides the sound, which may drift from it; `1.0` = the track exactly.
releaseINT80–64Audio steps the hold eases over where the track starts, ends or loops inside the video, at 40 a second: `0` = a hard edge; `8` = 0.2 s; `20` = half a second for the model to blend in and out.
past_endCOMBOmodel soundWhat plays once the track runs out before the video does: `model sound` = whatever the model makes; `silence` = held quiet; `loop` = the track starts again, the seam eased over `release`.

Outputs (2)

NameTypeDescription
windowLATENTThe window holding its slice of the track, for the sampler's latent input.
reportSTRINGWhich stretch of the track the window holds, how firmly, and under which frames of the finished video.