Nodes/comfyui-minimax-h3-audio-T8/H3 Avatar · HIGH Handoff — Recording Anchor (T8 EXP)
ComfyUI Node

H3 Avatar · HIGH Handoff — Recording Anchor (T8 EXP)

Handing H3 Avatar Off To HIGH Without Touching The Audio Anchor

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Avatar · HIGH Handoff — Recording Anchor (T8 EXP)
  • avatar_source
  • low_boundary
  • model
  • sampler
  • lifted_av
  • video_noise
  • high_source
  • high_sigmas
  • high_restart
  • plan
  • report_json

The Avatar route is a two-stage upsample: render LOW, lift the joint latent through the learned 3D latent upscaler, then run a HIGH pass to finish. The messy part is the seam between lift and HIGH - you have a latent that's been spatially blown up by a network that doesn't know anything about your recording, and you need the HIGH stage to continue without letting the audio anchor drift.

This node prepares that handoff. It does no sampling at all; it builds the HIGH restart state and hands you a plan and a report.

What it takes and what it restores

Inputs: avatar_source (the bound source from MiniMaxH3AvatarSourceBindEXPT8), low_boundary (the completed LOW stage's boundary), model, sampler, lifted_av (the upscaled latent), video_noise, and two optional ports - high_source and high_sigmas.

The docstring is explicit about the interesting constraint: by default the HIGH source is the bound recording source. You may supply a different high_source, but its clean audio and its zero audio mask must match the bound source. In other words, you can change the video side of the HIGH pass; you can't change what the audio anchor is. That's the rule that keeps a recording-driven avatar from turning into a voice-cloning pipeline by accident.

high_sigmas is optional too, for when you want a different remaining schedule on the HIGH pass. Because the audio clock isn't restarting, this is a partial refine, not a fresh generation.

Outputs: high_restart (typed T8_PROGRESSIVE_HIGH_RESTART), plan (typed T8_PROGRESSIVE_STAGE_PLAN) and report_json. The first two feed the common Progressive HIGH/EAV/Relay nodes - the same ones the non-Avatar progressive graphs use. That's deliberate: the Avatar entry points exist to bind the recording, not to fork the whole sampling stack.

Why there's a separate Avatar variant at all

The generic Progressive nodes don't know a recording is involved. The Avatar set - source bind, LOW, this handoff, and the delivery audit - threads the recording binding through every boundary so that the audio state stays locked to the real PCM rather than to whatever the model would have generated. If you've ever wondered why a pack needs four nodes that look like the generic ones, that's the reason.

Practical notes

  • lifted_av must be the output of the learned 3D latent upscaler run on the LOW boundary's video. Not the raw LOW output, not a VAE decode.
  • video_noise is the HIGH pass's noise. The pack's architecture separates the HIGH restart noise from the audio clean anchor on purpose, so don't reuse the LOW noise object here.
  • The first-期 scope is a single segment: no spatial tiling, no long-video chaining. The pack says so in the Avatar documentation, and it's worth not discovering that empirically.
  • Frame counts in this route run on H3's 17-frame grid; the pack's reference Avatar recipe uses 73 frames at 512×768, which is 5 + 17 × 4. Your own first-frame aspect ratio needs its own 32-pixel-aligned size, not a copy of that one.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Or ComfyUI Manager, search MiniMax H3 Audio T8. Fully quit ComfyUI and relaunch - no, a browser refresh doesn't count, node classes register at startup. Nothing to pip install; the pack declares no packages so it can't replace ComfyUI's Torch/CUDA stack. Needs recent ComfyUI with native H3 support. Weights are separate: H3 model, Qwen text encoder and the video/audio VAEs into models/diffusion_models, models/text_encoders and models/vae, plus the learned 3D latent upscaler wherever the pack's docs tell you to put it. There's no Avatar-specific checkpoint.

Where people trip

  • A high_source whose audio mask isn't zero. It's rejected, and rightly - a HIGH pass with an unlocked audio mask isn't the same computation.
  • Expecting this node to sample. It doesn't. Wire the outputs into the HIGH stage; if you queue just this, nothing renders.
  • Mixing generic Progressive boundaries with Avatar ones. They're typed differently on purpose. If a socket refuses, you've crossed the streams.
  • Reading a green run as a quality verdict. The pack is unusually clear that a passing Avatar recipe doesn't certify lip sync, voice or identity. Watch the clip.
CategoryT8/MiniMax H3/Modular Sampling/Avatar Experimental

Inputs (8)

NameTypeDefaultDescription
avatar_sourceT8_AVATAR_STAGE_SOURCE—
low_boundaryT8_PROGRESSIVE_LOW_BOUNDARY—
modelMODEL—
samplerSAMPLER—
lifted_avLATENT—
video_noiseNOISE—
high_sourceoptLATENT—
high_sigmasoptSIGMAS—

Outputs (3)

NameTypeDescription
high_restartT8_PROGRESSIVE_HIGH_RESTART—
planT8_PROGRESSIVE_STAGE_PLAN—
report_jsonSTRING—