H3 Audio Refine dual_clock · Bind Separate Tail (T8 EXP)
The Flexible End Of H3 Audio Refine
- plan
- original_av_latent
- model
- noise
- guider
- sampler
- sigmas
- stage_latent
- model
- noise
- guider
- sampler
- sigmas
- stage_latent
- stage_boundary
- report_json
If you only ever build one Audio Refine graph in this pack, it'll probably be this one. dual_clock is the flexible family: 1 to 8 refine steps, audio denoise anywhere between 0.01 and 1.0, and a seed you can reroll. The compatibility variant next to it in the node list is locked to 4 steps and a denoise of 0.35 or 0.50. Guess which one you want when you're still tuning.
What "dual clock" actually refers to
H3 denoises video and audio on separate sigma clocks inside one joint AV latent - the pack's reference recipe carries a video shift of 12 and an audio shift of 3. A tail refine is therefore an audio-clock operation: the video is frozen, the audio stream gets a handful of extra steps against it.
This node is the checkpoint at the front of that pass. You hand it the model, noise, guider, sampler and sigmas your dual-clock Setup produced, plus the frozen original_av_latent, the stage_latent, and the Setup's setup_report_json. It validates that the plan is signed, that the original AV carries the exact mask signature of an audio-only tail refine - 0 on video, 1 on audio - and that the report matches. Then it passes all five sampling inputs back out byte-identical and adds a stage_boundary.
It does not reschedule your noise, nudge your sigmas or swap your sampler. If you came here from packs where a "bind" node quietly rewires things, that's not what this is.
The sockets, in the order you'll touch them
Nine inputs, and they're all required - no optional ports, so a missing wire shows up immediately.
plan is typed H3_T8_AUDIO_REFINE_PLAN. That single type is the whole family discriminator: the dual_model bind wants H3_T8_AUDIO_REFINE_PHASE2_PLAN and the compat bind wants H3_T8_AUDIO_REFINE_COMPAT_PLAN. All three nodes read almost identically in the search box, so let the socket type do the work.
original_av_latent is the frozen first pass. stage_latent is the tensor the tail pass will actually sample. They aren't interchangeable and swapping them is the most common way to get an abort that reads like a mask problem.
model, noise, guider, sampler, sigmas, stage_latent go straight back out. stage_boundary is the token the matching MiniMaxH3AudioRefineDualClockStageAuditEXPT8 requires - there's no way to run the audit without it. report_json is your record; wire it somewhere visible.
One behaviour worth knowing up front: if the sigmas coming in are empty, this stays a no-sample path. Zero model calls, original AV passed through. That's the abstain branch, and the bind won't manufacture a schedule to hide it.
Installing
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Manager users can search MiniMax H3 Audio T8 instead - just be aware the GitHub release and the Registry listing move independently, so if you're chasing a specific version, git is the reliable route. Fully quit ComfyUI and relaunch afterwards; a browser refresh won't register anything new. No pip step here at all, by design: the pack's requirements file declares no packages so it can't replace ComfyUI's Torch or CUDA stack. You do need a recent ComfyUI with native H3 support, and the weights are separate downloads - H3 model in models/diffusion_models, Qwen in models/text_encoders, video and audio VAEs in models/vae.
Failure modes you'll actually meet
- A stale
setup_report_json. It's a receipt from the Setup run that produced the plan. Change a widget, re-run part of the graph, and validation fails - correctly. Re-run Setup. - Red plan wire. You grabbed the compat or dual_model plan. Different family, different node.
- A mask error on
original_av_latent. The node wants a genuine joint AV latent with audio locked, not a plain VAE-encoded latent. - Thinking this approves anything. It signs a boundary. The audit verifies, the Quality Gate decides, and default is to keep the original audio.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| plan | H3_T8_AUDIO_REFINE_PLAN | — | |
| original_av_latent | LATENT | — | |
| model | MODEL | — | |
| noise | NOISE | — | |
| guider | GUIDER | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| stage_latent | LATENT | — | |
| setup_report_json | STRING | — |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| noise | NOISE | — |
| guider | GUIDER | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| stage_latent | LATENT | — |
| stage_boundary | T8_AUDIO_REFINE_STAGE_BOUNDARY | — |
| report_json | STRING | — |