Nodes/comfyui-minimax-h3-audio-T8/FastH3 V2 · Separate LOW / HIGH Stage (T8 EXP)
ComfyUI Node

FastH3 V2 · Separate LOW / HIGH Stage (T8 EXP)

One DMD stage, no hidden loop behind the curtain

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
FastH3 V2 · Separate LOW / HIGH Stage (T8 EXP)
  • model
  • av_latent
  • model
  • sampler
  • sigmas
  • stage_context
  • report_json
◄stagelow_0_4►
◄profiletrained_vsa_exp►
◄min_tokens12288►

What it is, and why you'd reach for it

MiniMax H3 is the 33B omni-modal video model that MiniMax opened in August 2026 - video and native stereo audio in one pass, no separate audio stage bolted on. The catch, if you're in the US, EU, UK or Korea, is the licence: the open weights are geofenced and you're not licensed to run them locally at all. Everyone else gets a genuinely good model that is also slow, which is the entire reason FastH3 V2 exists.

FastH3 V2 is FastVideo's consistency-distilled 8-step student of H3 - a separate checkpoint, not a LoRA you stack on the base. Distillation always costs something (see distillation.md: every method out there trades quality for the step count), and the pack is honest about it. But 8 steps instead of the stock 20 is a large multiplier, and this node is how the pack lets you run those 8 steps yourself instead of inside one opaque mega-node.

The name is literal. This node sets up one DMD stage and never runs a hidden loop. Two of these, two SamplerCustomAdvanced nodes, and you have the whole chain: LOW 0:4 and HIGH 4:8, each with its own MODEL and LoRA stack.

How it works

It takes a MODEL and the AV latent, applies the FastH3 V2 stage profile, and hands back a patched MODEL plus the sampler/sigmas pair that correspond to that stage's slice of the schedule. LOW is the absolute 0:4 segment; HIGH is 4:8 with the AV shifts at 10/3. Nothing is sampled in here - report_json is the only side effect, and it's your receipt for what the stage actually built.

The important consequence: because the loop isn't hidden, every stage is independently replaceable. Cold-resume a HIGH-only run, swap a LoRA between passes, drop a stage save between them. That's the whole point of the modular rewrite.

The inputs that matter

  • model - your loaded FastH3 V2 ConvRot INT8 checkpoint, optionally through a LoRA or the pack's memory nodes.
  • av_latent - the nested joint AV latent coming out of the native H3 conditioning.
  • stage - low_0_4 (default) or high_4_8. Two instances of this node, one per stage.
  • profile - trained_vsa_exp by default, dense_compat_exp when you need the dense path. Read the next paragraph before you assume these are interchangeable.
  • min_tokens - the VSA sparse path's token floor, default 12288. The pack's accepted example recipes ran it at 0, so check which one you're on before comparing timings to anything you read.

Outputs: model (into the sampler), sampler and sigmas (into SamplerCustomAdvanced), stage_context (into the matching EAV and attestation nodes), and report_json.

The one trap worth knowing

Relay and trained VSA don't coexist yet. The trained sparse kernel has no per-query temporal bias adapter, so if you're using Prompt Relay on V2 you need the dense compatibility profile. Changing the profile dropdown and claiming you got both is not a thing. This bites people because the menu makes it look like a free choice.

Also worth repeating from the pack's own numbers: their cold/hot test on an RTX 4060 Ti 16GB showed V2 faster (~78s vs ~103s whole-graph) but with no VRAM reduction. A 22GB checkpoint is not 22GB of VRAM, and it is also not a promise of 16GB comfort.

Install

ComfyUI Manager → search MiniMax H3 Audio T8. Or manually:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Then fully quit and restart ComfyUI and refresh the browser - new nodes do not appear until you do. The pack's requirements.txt is deliberately empty (base nodes need only what ComfyUI already ships), so don't let anything pip-install over your Torch/CUDA stack. You need a recent ComfyUI with native H3 support, the H3 model in models/diffusion_models, the Qwen text encoder in models/text_encoders, and the video/audio VAEs in models/vae. For FastH3 V2 specifically, grab fastvideo_fasth3_8step_v2_pruned_int8_convrot.safetensors at the pinned revision the pack documents.

When it goes wrong

If the stage setup refuses to build, the usual cause is a mismatched input - a raw H3 MODEL where a V2 stage profile expects the ConvRot checkpoint, or a latent that isn't the nested AV shape. These EXP nodes fail loudly and early; read the traceback, it names the contract. And if ComfyUI Manager still shows an older version, that's normal: the GitHub release and the Registry listing are published independently, so install from GitHub when in doubt.

CategoryT8/MiniMax H3/Modular Sampling/Experimental

Inputs (5)

NameTypeDefaultDescription
modelMODEL—
av_latentLATENT—
stageCOMBOlow_0_42 options: low_0_4, high_4_8
profileCOMBOtrained_vsa_exp2 options: trained_vsa_exp, dense_compat_exp
min_tokensINT122880–1048576—

Outputs (5)

NameTypeDescription
modelMODEL—
samplerSAMPLER—
sigmasSIGMAS—
stage_contextT8_STAGE_CONTEXT—
report_jsonSTRING—