H3 Avatar · LOW Only — Recording Bound (T8 EXP)
H3 Avatar's Cheap Step, With The Recording Bound
- model
- sampler
- noise
- plan
- low_source
- positive
- negative
- avatar_source
- low_boundary
- report_json
Avatar work is slow, and most of the time you're iterating on things that don't need the high-resolution pass: is the head framing right, does the mouth open on the right syllable, did the first frame not look like a waxwork. This node runs the LOW stage of an H3 Avatar graph and nothing else. No hidden HIGH, no hidden lift.
What it does
It executes one initialised Avatar LOW pass and returns a boundary. That's the whole job, and the boundary is what makes it modular: low_boundary is the token the Avatar high handoff expects, and plan at the other end comes from the common Progressive stage plan running in initialized_av_exp mode. Feed the boundary into the lift and the HIGH handoff afterwards.
The Avatar-specific part is the last input. avatar_source - from MiniMaxH3AvatarSourceBindEXPT8 - carries the bound pair of encoded AV and original recording, with the audio mask pinned at zero. The LOW stage reads that binding so the audio state stays locked to your actual recording rather than to whatever the model felt like generating. There's no hidden HIGH, no lift, no second stage tucked inside; if the graph doesn't show it, it isn't happening.
Inputs worth knowing
model, sampler and noise are your LOW-stage sampling stack. plan must be the common Progressive plan in initialized_av_exp mode - the Avatar entry points don't accept a generic plan. low_source is the initial joint AV latent; positive and negative are conditioning.
cfg defaults to 1. That's not a typo or a conservative default, it's how H3 is normally run; the field goes up to 100 if you're doing something unusual, but leave it alone unless you have a reason.
reserve_vram_mib defaults to 1024 with a floor of 512. This is the resource guard: the stage checks the reserve at the start and again at each sampler callback, and it will surface a genuine out-of-memory condition rather than silently retrying with something different. Raise it if you're sharing a card with other work - the pack's own Avatar tests used a 6GiB reserve specifically to make ComfyUI offload more aggressively, and while that isn't a promise about your hardware, it's a useful hint that this knob does real work.
Outputs are low_boundary, typed T8_PROGRESSIVE_LOW_BOUNDARY, and report_json. There's no image output, because there's nothing to look at until HIGH finishes - or until you deliberately decode the LOW for a rough preview.
Where it sits
Progressive Avatar, end to end: source bind → LOW plan (initialized_av_exp) → a LOW conditioning node → this node → learned 3D lift → MiniMaxH3AvatarHighHandoffEXPT8 → HIGH conditioning → HIGH stage → MiniMaxH3AvatarDeliveryAuditEXPT8. EAV and Relay adapters attach independently to each stage if you're using them; the pack keeps them as separate external nodes rather than burying effects inside the stage node. Good design, slightly more wiring.
The LOW and HIGH stages want separate conditionings and separate noises. HIGH's conditioning dimensions should come from the actual upscaler output, not from a hand-typed number.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Or ComfyUI Manager → search MiniMax H3 Audio T8. Then quit ComfyUI fully and start it again - refreshing the page won't register new nodes. No packages to install: the pack's requirements file is empty by design so it can't replace your Torch/CUDA stack. You need a recent ComfyUI with native H3 support, and the weights are all separate - H3 diffusion model into models/diffusion_models, Qwen text encoder into models/text_encoders, video and audio VAEs into models/vae. Nothing Avatar-specific to download. Worth checking the H3 licence if you're in the US, EU, UK or South Korea; the weights are geofenced out of those territories.
Traps
- Queueing this expecting a video. It produces a boundary token. That's it. The decode happens after HIGH.
- Reusing the LOW conditioning for HIGH. Different sizes, different stage. The pack explicitly warns against feeding HIGH's dedicated conditioning into the LOW plan, because then changing a HIGH prompt re-runs LOW and your caching intuition quietly dies.
- A generic Progressive plan instead of
initialized_av_expmode. The Avatar stage wants the Avatar contract. - Expecting it to keep the recording as-is without the bind. If
avatar_sourceisn't carrying masks with audio locked, you're not running the Avatar computation.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| sampler | SAMPLER | — | |
| noise | NOISE | — | |
| plan | T8_PROGRESSIVE_STAGE_PLAN | — | |
| low_source | LATENT | — | |
| positive | CONDITIONING | — | |
| negative | CONDITIONING | — | |
| cfg | FLOAT | 1.00–100 | — |
| reserve_vram_mib | INT | 1024512–65536 | — |
| avatar_source | T8_AVATAR_STAGE_SOURCE | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| low_boundary | T8_PROGRESSIVE_LOW_BOUNDARY | — |
| report_json | STRING | — |