Nodes/comfyui-minimax-h3-audio-T8/H3 Avatar · Deliver Original Recording + Audit (T8 EXP)
ComfyUI Node

H3 Avatar · Deliver Original Recording + Audit (T8 EXP)

Your H3 Avatar Video Should Ship The Original Recording, Not A Model's Impression Of It

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Avatar · Deliver Original Recording + Audit (T8 EXP)
  • high_result
  • original_recording
  • av_latent
  • original_audio
  • report_json

Talking-head generation has a trap that doesn't show up until you've rendered a few: the model will happily produce audio for a face, and that audio is a plausible voice that isn't the recording you fed it. For an avatar - where the whole point is that a specific person said specific words - that's not a refinement, it's a defect.

This node is the pack's answer. It takes the finished HIGH result, verifies the recording binding that was established at the start of the graph, and returns the original recording PCM unchanged for your final mux, plus the audited AV latent and a report. It is not a lip-sync enhancer and it does not resynthesise the voice. It's the node that stops the graph from quietly replacing your recording with the model's idea of it.

How the binding works

Back at the front of the graph, MiniMaxH3AvatarSourceBindEXPT8 registered an encoded AV latent together with the original recording. The recording's audio is encoded into the latent with a zero audio mask - audio locked, video free - and that anchor is what the LOW and HIGH stages preserve. The delivery audit re-checks that anchor against the completed HIGH result before letting anything out.

It also accepts a saved HIGH result, which means you can resume and re-deliver a finished render without re-encoding the source or sampling anything. In a workflow where you've just burned an hour of GPU on a 73-frame clip, being able to re-audit and re-export without another pass is the difference between a usable iteration loop and none at all.

Inputs and outputs

Two inputs, both required: high_result, typed T8_PROGRESSIVE_HIGH_RESULT and produced by the Avatar/Progressive HIGH side of the graph, and original_recording typed AUDIO. That second socket takes the same recording object you bound at the start. Don't re-decode it, don't convert it, don't pass the model's generated audio - it's the reference the audit is checking against.

Three outputs. av_latent is the audited result, ready for your decode/export path. original_audio is the original PCM, unchanged, which is what you mux with the final video. report_json is the receipt, and since this is an output node it also renders on the node face.

The report includes the latent anchor difference - the measured gap between the delivered AV's audio anchor and the bound recording's. That number is diagnostic. It is not a quality score, and the node's own description says plainly that it does not certify trained lip sync, voice identity or overall quality. Those stay a human judgement. Believe it.

Where it fits

Avatar-wide, the pack's route is: first-frame image plus recording encoded through the audio conditioning node with audio_mode=lock_source and the recording on drive_audio, then the Avatar LOW stage, then the learned 3D lift, then the HIGH handoff and HIGH stage, then this delivery audit. There's no Avatar-specific model to download; it uses the same native H3 set plus your chosen learned 3D latent upscaler.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Or search MiniMax H3 Audio T8 in ComfyUI Manager. Fully quit ComfyUI and relaunch afterwards; a browser refresh won't pick up new nodes. There's no pip step, deliberately - the pack ships an empty requirements file so installing it can never replace the Torch or CUDA build your ComfyUI is using. You will need a recent ComfyUI with native H3 support, and weights are separate: H3 in models/diffusion_models, Qwen in models/text_encoders, video and audio VAEs in models/vae. One licensing note if you're in the US, EU, UK or South Korea: the H3 weights are geofenced out of those territories by the MiniMax H3 Community License, so check that before you invest a week in a pipeline.

What goes wrong

  • Passing generated audio as original_recording. The audit is checking that you didn't. It will notice, and you've defeated the node's purpose anyway.
  • Reading the anchor difference as a pass/fail. It's a measurement, not a verdict. The verdict is you watching the clip.
  • Forgetting it's an output node. It has no UI switch to enable; queuing it delivers.
  • Expecting it to fix lip sync. It doesn't touch picture, and the pack is explicit that nothing here certifies lip sync or voice identity.
CategoryT8/MiniMax H3/Modular Sampling/Avatar Experimental

Inputs (2)

NameTypeDefaultDescription
high_resultT8_PROGRESSIVE_HIGH_RESULT—
original_recordingAUDIO—

Outputs (3)

NameTypeDescription
av_latentLATENT—
original_audioAUDIO—
report_jsonSTRING—