Nodes/comfyui-minimax-h3-audio-T8/H3 Avatar 录音驱动渐进采样(EXP/T8)
ComfyUI Node

H3 Avatar 录音驱动渐进采样(EXP/T8)

Make a portrait talk with a recording you already have — no cloning, no retraining

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
H3 Avatar 录音驱动渐进采样(EXP/T8)
  • model
  • positive
  • negative
  • av_latent
  • sampler
  • sigmas
  • model_hires
  • av_latent
  • report_json
◄upscaler_model▾►
◄seed20260909►
◄cfg1.0►
◄low_evaluations4►
◄low_scale0.50►
◄taski2va►
◄precisionfp16►
◄reserve_vram_mib1024►
◄guide_resizelegacy_bilinear►
◄eav_modedisabled►
◄eav_tau4.0►
◄eav_start_video_progress0.15►
◄eav_end_video_progress0.90►
◄eav_max_workspace_mib32►
◄eav_g_hard_limit1.50►

You have a still portrait and a voice recording, and you want the first and not a new performance of the second. That's the exact case MiniMaxH3AvatarProgressiveEXPT8 exists for: the recording is encoded into the H3 audio latent and held there with a zero audio mask through both sampling stages, so what ships in the file is your original track driving the mouth - not a generated line in a similar timbre.

Two things it is not. It's not voice cloning: it will not say new sentences in that voice. And it's not an audio mux - the recording participates in sampling, then is delivered as-is.

The shape of the sampling

It's a two-stage progressive run: sample the LOW stage at a reduced low_scale, push the latent through the learned 3D H3 upscaler, then finish the remaining steps at full size with the HIGH stage. Default is 4 + 4, and low_evaluations is how many evaluations belong to LOW - the author's tooltip gives the arithmetic: in a full 8-step schedule, pick 6 for the small size and you get 2 more steps after upscaling. At least 1 step has to remain.

low_scale (0.5) is the small-stage width/height ratio, snapped to 32-pixel alignment. The tooltip is refreshingly honest here: 0.5 does not guarantee a 4x speedup. The pack's own docs repeat it - no universal speed or VRAM claim for progressive sampling.

The inputs that actually matter

Wire the same clock the native DualClock Setup uses. The sampler tooltip is specific: connect the native Setup's euler, and don't use dual_clock_euler here. sigmas wants the complete first-pass schedule, from 1 to 0, at least 2 steps.

av_latent is the native nested H3 audio/video latent, and it has to come from the upstream MiniMaxH3AudioConditioningT8 node set to I2VA with audio_mode=lock_source and your recording on drive_audio. That's where the "your voice, not a clone" contract is established; this node inherits it.

upscaler_model must be picked explicitly - the H3 learned latent upscaler in models/latent_upscale_models. precision (fp16), cfg (1.0) and seed are what you'd expect. task defaults to i2va, and reserve_vram_mib (1024) is a headroom check at stage boundaries and sampling callbacks - the tooltip is clear it is not a peak prediction and won't stop an OOM.

Optional, and both are opt-in for a reason:

  • model_hires lets the HIGH stage use an independent MODEL, so you can chain different LoRA or attention patches per stage. Leave it unconnected and HIGH reuses model.
  • guide_resize (legacy_bilinear or preserve_mean) controls how the first-frame LOW reference is scaled. preserve_mean explicitly holds channel means; the HIGH reference is untouched either way.
  • The eav_* family is per-stage Enhance-A-Video. eav_mode defaults to disabled, with report_only to measure and apply_exp to actually apply. eav_tau = 0 is not a means of turning it off; disabled is.

Outputs are av_latent and report_json. Decode the latent with the normal H3 VAE; read the report when you want to prove which policies were in force.

Installing

ComfyUI Manager → MiniMax H3 Audio T8, or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Full restart of ComfyUI, then refresh the page. No extra packages: the pack's requirements.txt deliberately installs nothing so it can't stomp your Torch/CUDA stack. You do need a recent Core with native H3 support, the H3 base in models/diffusion_models, Qwen3-VL in models/text_encoders, video and audio VAEs in models/vae, the learned 3D upscaler in models/latent_upscale_models, and any Avatar EMA-B accelerator LoRA in models/loras. There is no Avatar-specific checkpoint to download - Avatar reuses what you already have.

Where it gets fiddly

Single segment only. The first phase of this entry point is one standalone clip; there's no spatial tiling and no long-video qualification. If you want 8 seconds across two segments, that's the pack's other workflows, not this node.

Match the AudioWindow to the clip. Set duration to frames/24 with warmup and cooldown at 0 and ensure_minimum_context=false, or the conditioning will quietly expand a short test into a longer window. The worked example is 73 frames at 24 fps ≈ 3.04 s.

Recording length has to cover the clip. Short input is zero-padded, long input is windowed - but if the start of your track is a breath, don't also stack a big opening mute and fade on top of it, or you'll eat the first word.

Judge it by watching and listening. Numeric checks stay numeric checks: lip sync and delivery still need your eyes and ears, and other recipes' seam results don't transfer here.

CategoryT8/MiniMax H3/Audio/Experimental

Inputs (22)

NameTypeDefaultDescription
modelMODEL—
positiveCONDITIONING—
negativeCONDITIONING—
av_latentLATENT—
samplerSAMPLER连接原生 Setup 的 euler,不使用 dual_clock_euler。
sigmasSIGMAS完整首采时间表:从 1 到 0,至少 2 步。
upscaler_modelCOMBOmodels/latent_upscale_models 内的 H3 学习型潜空间放大模型。必须显式选择。
seedINT202609090–18446744073709550000—
cfgFLOAT1.00–100—
low_evaluationsINT41–999小画幅的步数。例:完整 8 步中选 6,则放大后还采 2 步。至少留 1 步。
low_scaleFLOAT0.500.25–0.99小画幅宽高比例,按 32 像素对齐;0.5 不是保证 4 倍加速。
taskCOMBOi2va2 options: t2va, i2va
precisionCOMBOfp163 options: fp16, bf16, fp32
reserve_vram_mibINT1024512–32768阶段边界和采样回调检查的显存余量,不代表峰值预测或不会 OOM。
model_hiresoptMODEL可选HIGH阶段MODEL;可独立串联LoRA/注意力补丁,原有补丁保留。结构和AV时钟须匹配;不接沿用model。
guide_resizeoptCOMBOlegacy_bilinear首帧LOW参考缩放;preserve_mean显式保持通道均值,HIGH原参考不变。
eav_modeoptCOMBOdisabled分阶段EAV使用完整8/20步原视频sigma时钟,CFG1。report_only只测量,apply_exp实际增强;不是画质认证。
eav_tauoptFLOAT4.0-32–32—
eav_start_video_progressoptFLOAT0.150–11-原视频sigma;切换分辨率不重置。旧API省略该参数仍保留原0/1合同。
eav_end_video_progressoptFLOAT0.900–1—
eav_max_workspace_miboptINT324–512—
eav_g_hard_limitoptFLOAT1.501–3真实超限正常报错,不截断或重试。关闭选disabled,tau=0不是关闭。

Outputs (2)

NameTypeDescription
av_latentLATENT—
report_jsonSTRING—