Nodes/comfyui-minimax-h3-audio-T8/H3 Audio Refine · Audit Sample or Abstain (T8 EXP)
ComfyUI Node

H3 Audio Refine · Audit Sample or Abstain (T8 EXP)

The Last Automated Gate Before Your Refined H3 Audio Tail Ships

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Audio Refine · Audit Sample or Abstain (T8 EXP)
  • second_pass_input
  • second_pass_output
  • stage_boundary
  • verified_av_latent
  • report_json
◄expected_audio_strength1.00►
◄fail_on_locked_mismatchfalse►
◄locked_atol0.00►
◄stage_report_json—►

Audio Refine ends with an uncomfortable question: did the tail pass actually do what the plan said it would - touch the audio, leave the video alone - or did it quietly do something adjacent? This node is where the pack answers that, and it's honest about the fact that sometimes the answer is "nothing happened, on purpose".

Two behaviours, one node

If the stage boundary says a tail was sampled, this runs the pack's original strict two-pass audio audit unchanged: it compares the state before the second pass against the state after, and checks the audio where it should have moved against the video where it should not.

If the boundary says the plan signed ABSTAIN, it requires the original AV exactly as it came in and passes it straight through. Critically, it does not invent a sampling mask and does not claim an accepted candidate just to produce a tidy output. A no-sample run stays a no-sample run all the way to delivery, and the report says so.

The knobs that matter

second_pass_input is the AV latent going into the second pass; second_pass_output is what came out. Yes, that means you wire the same tensor to two places - that's the point, and swapping them will produce a confidently wrong report rather than a crash.

stage_boundary and stage_report_json come from the stage bind pair upstream and pin this audit to one specific pass. Without them it has no idea what it's supposed to be auditing.

expected_audio_strength is a FLOAT from 0 to 1, defaulting to 1. That's your declared strength for the audio change - the audit measures against it, so if you set this to something arbitrary the report stops meaning anything. Leave it at 1 unless your plan genuinely says otherwise.

Then the strictness pair. locked_atol is the tolerance for the region that must not have changed - the video side of the joint latent - and it defaults to 0, meaning exact. fail_on_locked_mismatch defaults to false. That combination is the sensible starting point: you get the report, including any locked-region drift, without the graph dying on your first attempt. Once you trust the wiring, flipping fail_on_locked_mismatch to true turns the same check into a hard stop. Keep the default zero tolerance; raising locked_atol to make a warning disappear is how you turn an audit into decoration.

Outputs are verified_av_latent and report_json, and it's an output node, so the report is drawn on the node itself. Read it. It's the only place you learn whether you got a refined tail or a passthrough.

Where it sits

Downstream of the stage audit, upstream of the Quality Gate. It's the last thing that happens without a human involved, and it deliberately stops short of approving anything - the pack is blunt that the gate and human listening remain authoritative, and that the default delivery choice keeps the original audio. If you're expecting this node to select the refined take for you, no node in this pack does that.

Install and requirements

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Manager equivalent: search MiniMax H3 Audio T8. Then fully quit and restart ComfyUI - a page refresh won't do it. The pack's requirements file contains no packages at all by design, so there's no dependency install and, more importantly, no chance of it swapping out the Torch or CUDA build ComfyUI is running on. What you do need: a recent ComfyUI with native H3 support, and your own weights - H3 in models/diffusion_models, Qwen text encoder in models/text_encoders, and the video and audio VAEs in models/vae.

Snags worth knowing

  • Reading a passthrough as a failure. An ABSTAIN tail passing the original AV through is the designed outcome, not a bug.
  • expected_audio_strength left at a guessed value. The report's numbers become meaningless against a strength you didn't actually use.
  • Hard-failing on run one because you flipped fail_on_locked_mismatch before you'd seen a clean report. Start permissive, read, then tighten.
  • Treating a pass as a quality verdict. It verifies shape and drift, not whether the audio sounds better. That's still your job.
CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (7)

NameTypeDefaultDescription
second_pass_inputLATENT—
second_pass_outputLATENT—
expected_audio_strengthFLOAT1.000–1—
fail_on_locked_mismatchBOOLEANfalse—
locked_atolFLOAT0.000–1—
stage_boundaryT8_AUDIO_REFINE_STAGE_BOUNDARY—
stage_report_jsonSTRING—

Outputs (2)

NameTypeDescription
verified_av_latentLATENT—
report_jsonSTRING—