Nodes/comfyui-minimax-h3-audio-T8/H3 Avatar · Bind Encoded Source + Recording (T8 EXP)
ComfyUI Node

H3 Avatar · Bind Encoded Source + Recording (T8 EXP)

H3 Avatar's Smallest And Most Important Node

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Avatar · Bind Encoded Source + Recording (T8 EXP)
  • high_source
  • original_recording
  • high_source
  • avatar_source
  • report_json

Two inputs, three outputs, no computation. This node registers the pair your whole Avatar render depends on: the encoded AV latent and the original recording, bound together as a T8_AVATAR_STAGE_SOURCE that every downstream Avatar stage reads from.

It's the least impressive node in the Avatar set and the one that determines whether the rest means anything.

Why it exists

H3 is a joint audio-video model, so a talking-head render is generated as one latent with the audio inside it. If you let that run free, the audio you get is plausible and isn't your recording. For an avatar that's the wrong answer. So the pack's Avatar route locks the audio: the recording's audio is encoded into the latent with a mask of zero, meaning "known, don't generate", while the video side stays free. The pipeline carries that locked state through LOW, the learned lift, and HIGH.

This node is where the lock is established and, more importantly, where it's recorded - so the delivery audit at the end can verify the finished render is still anchored to the pair you chose at the start.

It records, it doesn't convert

high_source is your encoded AV latent. original_recording is the AUDIO object - the actual PCM you want in the final mux. The node encodes nothing and samples nothing. Its own description is careful about the distinction: it records your selected pair, and that isn't the same as proving where the latent came from. If your latent came out of some other pipeline, the binding will faithfully remember that you chose it.

That honesty matters later. When MiniMaxH3AvatarDeliveryAuditEXPT8 reports a latent anchor difference, it's measuring against this recorded pair, not against an assumed provenance.

The mask requirement is the real gate

Here's the thing that will actually stop you: the input latent needs explicit nested AV masks with audio=0. The docstring points at the pack's Audio Latent Control lock as the way to produce that. If you hand it a plain latent from a normal encode, there's no nested mask structure to inspect and it won't bind.

If you're new to this: in this pack's convention, a mask of 0 means "known / locked" and 1 means "generate". An audio mask of all zeros is what makes this a recording-driven avatar instead of a sound-alike generator. That one convention is the difference between the two things, and it's worth being able to say out loud before you debug anything.

Outputs

high_source is your latent passed straight through, so you can keep your existing wiring downstream. avatar_source is the typed T8_AVATAR_STAGE_SOURCE - this goes into the LOW stage, the high handoff, and anything else Avatar-typed. report_json records what got bound, and you should glance at it once, because it tells you what the pipeline thinks it's working with.

One instruction worth following literally: keep this source independent of HIGH prompts. The node's description says so, and the reason is graph dependency - if your neutral source depends on HIGH conditioning, every HIGH prompt change invalidates the LOW work you were trying to keep.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Manager also works - search MiniMax H3 Audio T8 - but the author notes GitHub and Registry releases move independently, so clone when you care about versions. Fully quit ComfyUI and relaunch. No pip step: the pack's requirements file declares no packages at all, so it can't disturb ComfyUI's Torch or CUDA build. Requires a recent ComfyUI with native H3 support. Weights are separate and none are Avatar-specific: the H3 model goes in models/diffusion_models, Qwen in models/text_encoders, video and audio VAEs in models/vae.

What goes wrong

  • A latent without nested AV masks. The most common failure, and it happens before any GPU work. Build the source through the audio-latent lock path.
  • Passing generated audio as the recording. The bind will accept it, and the audit will later tell you the anchor doesn't match your expectation. Use the real file.
  • Expecting validation of provenance. It isn't a VAE claim. It remembers your selection.
  • Letting a HIGH prompt change invalidate the source. Keep the binding neutral and independent, as the description says.
CategoryT8/MiniMax H3/Modular Sampling/Avatar Experimental

Inputs (2)

NameTypeDefaultDescription
high_sourceLATENT—
original_recordingAUDIO—

Outputs (3)

NameTypeDescription
high_sourceLATENT—
avatar_sourceT8_AVATAR_STAGE_SOURCE—
report_jsonSTRING—