FastH3 V2 · 8-Step Recipe (T8 EXP)
FastH3 V2 is not a LoRA — it's the 8-step recipe node for a different student model
- model
- av_latent
- model
- sampler
- sigmas
- report_json
What this node actually is
The name undersells it. "8-Step Recipe" sounds like a sampler preset; what you actually plug in is a MODEL plus an audio-video latent, and what comes back is a patched MODEL, a SAMPLER, a SIGMAS, and a JSON report. The author's own description says the important part out loud: this is a "Full V2 student, not a four-step LoRA."
That distinction is the whole point. The ComfyUI reflex is to slap an 8-step Turbo-style LoRA on the base model and declare it fast. FastH3 V2 doesn't work that way. The speed comes from a separately trained student checkpoint - fastvideo_fasth3_8step_v2_pruned_int8_convrot.safetensors, ~22.1 GB, from the FastVideo project, with the README pinning both the revision and a SHA256 - plus the exact 8-rung denoising ladder that student was trained for. Our distillation write-up covers the general trade: a student collapses a long trajectory into a few big jumps, and every method of doing that gives up some quality to get there. This node is the plumbing that makes the trained ladder actually fire in ComfyUI.
How it works
The ladder is rungs 999, 874, 749, 624, 500, 375, 250, 125 over 1000 with a final 0 - eight real forwards, not a uniform nine-point schedule. Video and audio run on separate clocks, shift 10 and shift 3 respectively, each computed as s*q/(1+(s-1)*q). Then the node installs an attention owner: eligible block attention goes through H3's gate-aware chunked sparse path (learned gates, 64-token tiles, ~20% kept).
It validates hard before it does any of that. The model has to be a native MiniMaxH3Model - an OpenVDN-architecture or plain base H3 checkpoint is rejected by name - every block has to carry an intact learned gate of the expected shape, head_dim has to be 128, and Core's comfy_extras.nodes_sparse_attention has to expose the sparse/override/eligibility interfaces it calls. If your ComfyUI predates those, you get "requires native H3 gate-aware chunked VSA support" rather than a silent bad video. Update ComfyUI first; that's the fix.
The three profiles are not the same recipe
trained_vsa_exp (default) is the trained table plus learned VSA. dense_compat_exp keeps the same DMD table and init but preserves whatever Dense backend you've put upstream - a KJ Sage or Sol selector - and it's the only profile that can carry Relay's timeline bias. official_comfy_template_exp switches to Core simple8/res_multistep with 10% keep and a 20% start: a comparison recipe lifted from the official Comfy template, not the trained integration. The author is blunt that these can't be called equivalent.
Inputs and outputs that matter
Four inputs do the real work. model takes the loaded student, optionally through a plain-weight LoRA or the T8 LowVRAM/ChunkFFN nodes - but not through an old EMA-B or Turbo speed LoRA, which will fight the recipe. av_latent is the native H3 AV conditioning output; it isn't decorative, the node reads its shape to know the packed token counts and to set the audio clock. profile is your pick from above. min_tokens (default 12288) is the eligibility floor: below that many packed tokens, native VSA runs Dense and admits it in the report; setting 0 forces eligibility testing, which is what the small probe workflows do. It's a performance qualifier, not a VRAM gate.
On the way out, model goes to your guider's model input, sampler and sigmas to SamplerCustomAdvanced, and BasicGuider runs at CFG 1 with your usual RandomNoise. Decode the sampler's zero-sigma output - not a mid-run x0 preview. report_json is a string; wire it into a text node if you want to read it, or use the paired audit node.
Installing it
ComfyUI Manager, search MiniMax H3 Audio T8, then a full restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
No extra pip packages - the pack's requirements.txt is deliberately empty so it can't touch your Torch/CUDA stack. You supply the FastH3 student in models/diffusion_models/, and it reuses your existing Qwen3-VL text encoder and the native video/audio VAEs. Node-all-red after install means your ComfyUI itself is too old, which is the same advice the ecosystem doc gives for missing-node pileups generally.
What to expect, and one licence warning
The only number the author published: 832×480, 73 frames, RTX 4060 Ti 16GB - trained DMD8 at 78.20s cold / 71.92s hot versus the old EMA-B native 8-step at 103.35 / 99.90. Roughly 24–28% faster wall clock, and cycle-max VRAM did not drop. So this is a speed path, not a low-VRAM path, and it's one machine and one prompt, not a quality-equivalence proof.
Last thing: MiniMax H3's weights sit under the geofenced community licence that excludes the EU, UK, Korea and the US, and derived H3 models follow it (panel). Worth thirty seconds of reading before pulling 22 GB.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| av_latent | LATENT | — | |
| profile | COMBO | trained_vsa_exp | 3 options: trained_vsa_exp, dense_compat_exp, official_comfy_template_exp |
| min_tokens | INT | 122880–1048576 | Below this packed-token count native VSA runs Dense and reports it; 0 forces eligibility testing. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| report_json | STRING | — |