H3 Avatar 录音驱动渐进采样(EXP/T8)
Make a portrait talk with a recording you already have — no cloning, no retraining
- model
- positive
- negative
- av_latent
- sampler
- sigmas
- model_hires
- av_latent
- report_json
You have a still portrait and a voice recording, and you want the first and not a new performance of the second. That's the exact case MiniMaxH3AvatarProgressiveEXPT8 exists for: the recording is encoded into the H3 audio latent and held there with a zero audio mask through both sampling stages, so what ships in the file is your original track driving the mouth - not a generated line in a similar timbre.
Two things it is not. It's not voice cloning: it will not say new sentences in that voice. And it's not an audio mux - the recording participates in sampling, then is delivered as-is.
The shape of the sampling
It's a two-stage progressive run: sample the LOW stage at a reduced low_scale, push the latent through the learned 3D H3 upscaler, then finish the remaining steps at full size with the HIGH stage. Default is 4 + 4, and low_evaluations is how many evaluations belong to LOW - the author's tooltip gives the arithmetic: in a full 8-step schedule, pick 6 for the small size and you get 2 more steps after upscaling. At least 1 step has to remain.
low_scale (0.5) is the small-stage width/height ratio, snapped to 32-pixel alignment. The tooltip is refreshingly honest here: 0.5 does not guarantee a 4x speedup. The pack's own docs repeat it - no universal speed or VRAM claim for progressive sampling.
The inputs that actually matter
Wire the same clock the native DualClock Setup uses. The sampler tooltip is specific: connect the native Setup's euler, and don't use dual_clock_euler here. sigmas wants the complete first-pass schedule, from 1 to 0, at least 2 steps.
av_latent is the native nested H3 audio/video latent, and it has to come from the upstream MiniMaxH3AudioConditioningT8 node set to I2VA with audio_mode=lock_source and your recording on drive_audio. That's where the "your voice, not a clone" contract is established; this node inherits it.
upscaler_model must be picked explicitly - the H3 learned latent upscaler in models/latent_upscale_models. precision (fp16), cfg (1.0) and seed are what you'd expect. task defaults to i2va, and reserve_vram_mib (1024) is a headroom check at stage boundaries and sampling callbacks - the tooltip is clear it is not a peak prediction and won't stop an OOM.
Optional, and both are opt-in for a reason:
model_hireslets the HIGH stage use an independent MODEL, so you can chain different LoRA or attention patches per stage. Leave it unconnected and HIGH reusesmodel.guide_resize(legacy_bilinearorpreserve_mean) controls how the first-frame LOW reference is scaled.preserve_meanexplicitly holds channel means; the HIGH reference is untouched either way.- The
eav_*family is per-stage Enhance-A-Video.eav_modedefaults todisabled, withreport_onlyto measure andapply_expto actually apply.eav_tau= 0 is not a means of turning it off;disabledis.
Outputs are av_latent and report_json. Decode the latent with the normal H3 VAE; read the report when you want to prove which policies were in force.
Installing
ComfyUI Manager → MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Full restart of ComfyUI, then refresh the page. No extra packages: the pack's requirements.txt deliberately installs nothing so it can't stomp your Torch/CUDA stack. You do need a recent Core with native H3 support, the H3 base in models/diffusion_models, Qwen3-VL in models/text_encoders, video and audio VAEs in models/vae, the learned 3D upscaler in models/latent_upscale_models, and any Avatar EMA-B accelerator LoRA in models/loras. There is no Avatar-specific checkpoint to download - Avatar reuses what you already have.
Where it gets fiddly
Single segment only. The first phase of this entry point is one standalone clip; there's no spatial tiling and no long-video qualification. If you want 8 seconds across two segments, that's the pack's other workflows, not this node.
Match the AudioWindow to the clip. Set duration to frames/24 with warmup and cooldown at 0 and ensure_minimum_context=false, or the conditioning will quietly expand a short test into a longer window. The worked example is 73 frames at 24 fps ≈ 3.04 s.
Recording length has to cover the clip. Short input is zero-padded, long input is windowed - but if the start of your track is a breath, don't also stack a big opening mute and fade on top of it, or you'll eat the first word.
Judge it by watching and listening. Numeric checks stay numeric checks: lip sync and delivery still need your eyes and ears, and other recipes' seam results don't transfer here.
Inputs (22)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| positive | CONDITIONING | — | |
| negative | CONDITIONING | — | |
| av_latent | LATENT | — | |
| sampler | SAMPLER | 连接原生 Setup 的 euler,不使用 dual_clock_euler。 | |
| sigmas | SIGMAS | 完整首采时间表:从 1 到 0,至少 2 步。 | |
| upscaler_model | COMBO | models/latent_upscale_models 内的 H3 学习型潜空间放大模型。必须显式选择。 | |
| seed | INT | 202609090–18446744073709550000 | — |
| cfg | FLOAT | 1.00–100 | — |
| low_evaluations | INT | 41–999 | 小画幅的步数。例:完整 8 步中选 6,则放大后还采 2 步。至少留 1 步。 |
| low_scale | FLOAT | 0.500.25–0.99 | 小画幅宽高比例,按 32 像素对齐;0.5 不是保证 4 倍加速。 |
| task | COMBO | i2va | 2 options: t2va, i2va |
| precision | COMBO | fp16 | 3 options: fp16, bf16, fp32 |
| reserve_vram_mib | INT | 1024512–32768 | 阶段边界和采样回调检查的显存余量,不代表峰值预测或不会 OOM。 |
| model_hiresopt | MODEL | 可选HIGH阶段MODEL;可独立串联LoRA/注意力补丁,原有补丁保留。结构和AV时钟须匹配;不接沿用model。 | |
| guide_resizeopt | COMBO | legacy_bilinear | 首帧LOW参考缩放;preserve_mean显式保持通道均值,HIGH原参考不变。 |
| eav_modeopt | COMBO | disabled | 分阶段EAV使用完整8/20步原视频sigma时钟,CFG1。report_only只测量,apply_exp实际增强;不是画质认证。 |
| eav_tauopt | FLOAT | 4.0-32–32 | — |
| eav_start_video_progressopt | FLOAT | 0.150–1 | 1-原视频sigma;切换分辨率不重置。旧API省略该参数仍保留原0/1合同。 |
| eav_end_video_progressopt | FLOAT | 0.900–1 | — |
| eav_max_workspace_mibopt | INT | 324–512 | — |
| eav_g_hard_limitopt | FLOAT | 1.501–3 | 真实超限正常报错,不截断或重试。关闭选disabled,tau=0不是关闭。 |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| av_latent | LATENT | — |
| report_json | STRING | — |