Nodes/comfyui-minimax-h3-audio-T8/H3 Audio Refine · Audit Separate Tail Candidate (T8 EXP)
ComfyUI Node

H3 Audio Refine · Audit Separate Tail Candidate (T8 EXP)

This H3 Audio Refine Audit Refuses To Approve Your Candidate — That's The Point

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Audio Refine · Audit Separate Tail Candidate (T8 EXP)
  • stage_boundary
  • plan
  • original_av_latent
  • stage_latent
  • candidate_av_latent
  • candidate_av_latent
  • report_json
◄setup_report_json—►

MiniMax H3 doesn't generate a picture and then dub sound over it. It generates one joint AV latent, with stereo audio baked into the same denoising run. That's why "make the audio better" in this pack means "resample part of that joint latent" - and why the pack's Audio Refine family exists at all: after your video is frozen, you spend a few extra steps refining only the audio tail.

This node is the second half of the compatibility tail-refine pair. It runs after your external sampler has produced a candidate, rechecks the source and the signed stage boundary, and hands the candidate back with a report. What it deliberately does not do is accept anything.

What it actually checks

The pack's modular sampling philosophy is that a graph connection isn't evidence. So the audit re-derives, from the same inputs, what the bind node asserted earlier: that original_av_latent is the original AV (not the candidate), that stage_latent is the one the stage was bound to, and that setup_report_json still matches the setup that signed this run. If any of that drifted, you get an error rather than a plausible-looking output file.

That report is a receipt, not a setting. It's produced by the Setup side of the graph, so if you changed a widget upstream and re-ran only part of the graph, the mismatch is your answer.

The important behaviour is the abstain path. When the plan signs ABSTAIN, the audit returns the original AV unchanged - no invented sampling mask, no "well, the candidate looks close" fallback. Empty SIGMAS upstream means no sampling happened at all, and the audit keeps that honest rather than papering over it.

Inputs and outputs you touch

Your stage's plan, stage_boundary, original_av_latent, stage_latent and setup_report_json all come from the matching MiniMaxH3AudioRefineCompatStageBindEXPT8 outputs - you're not typing any of them. candidate_av_latent is the sampler result you want checked.

It hands back candidate_av_latent and report_json. It's an output node, so the report prints on the node face too. Read the report; it's the only place the audit tells you why it abstained or what the residual difference was.

Note the plan socket is typed H3_T8_AUDIO_REFINE_COMPAT_PLAN. The other two variants of this node in the pack - the dual_clock and dual_model audits - share this node's exact display name, but their plan sockets are different types. ComfyUI will refuse the wire, which is the one mercy in this naming scheme.

Why the compatibility flavour is the strict one

The three Audio Refine plan families are not interchangeable presets. The compatibility plan, which feeds this node, is the most locked-down: the source enforces exactly 4 refine steps and an audio denoise of either 0.35 or 0.50, nothing else. The dual_clock plan is the flexible one (1–8 steps, denoise anywhere from 0.01 to 1.0). If you were hoping to try 0.6 here, that's not this node being broken - it's the compatibility contract.

Installing it

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Manager works too - search MiniMax H3 Audio T8. Then fully quit ComfyUI and restart; a page refresh alone won't re-register nodes. There's no requirements.txt to fight with: the pack deliberately declares zero extra packages so it can never clobber your Torch/CUDA stack. It does need a recent ComfyUI with native H3 support, and it ships no weights - the H3 model, Qwen text encoder and video/audio VAEs go in models/diffusion_models, models/text_encoders and models/vae respectively.

Where people get burned

  • Running this audit but never the Quality Gate. The audit reports; the gate decides. The pack's default is to keep the original audio, and human listening stays authoritative.
  • Reusing a setup_report_json from an earlier setup run. It won't validate, and you'll swear the node is buggy.
  • Expecting the candidate to be "applied". It isn't, here. This is the checkpoint, not the turnstile.
  • Two copies of this pack in custom_nodes. It has bitten people before - an old test copy can shadow the real one and hand you a stack trace from a completely different Python module.
CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (6)

NameTypeDefaultDescription
stage_boundaryT8_AUDIO_REFINE_STAGE_BOUNDARY—
planH3_T8_AUDIO_REFINE_COMPAT_PLAN—
original_av_latentLATENT—
stage_latentLATENT—
setup_report_jsonSTRING—
candidate_av_latentLATENT—

Outputs (2)

NameTypeDescription
candidate_av_latentLATENT—
report_jsonSTRING—