H3 VDN · Complete / Own-Grid Tail Stage (T8 EXP)
Setting up one VDN stage without hidden sampling
- model
- av_latent
- first_pass_latent
- model
- sampler
- sigmas
- stage_context
- report_json
MiniMax H3 at stock step counts is a lot of generation for a hobby card - the announcement threads were full of people asking what "slowly" meant on a 3060, and the honest answer is "get comfortable". OpenVDN is the community answer: a distilled H3 route that gets a usable clip out in eight DMD steps, plus a 50-step Stage B variant if you'd rather have quality. It's the pack's recommended acceleration route, and this node is how you drive it one stage at a time.
MiniMaxH3VDNStageSetupEXPT8 prepares a single VDN stage - the full distilled trajectory, or its own fresh-noise tail - and hands you the sampler and sigmas. It loads no weights and samples nothing.
Inputs
model- an OpenVDN Composer model, and this is a real requirement, not a suggestion. The node checks the model's own configuration receipt: right schema, status configured, a known stage, and step count matching. Hand it a plain H3 model and it errors instead of improvising.av_latent- the stage's AV canvas.stage-vdn_complete(the full DMD8/B50 first pass, the default) orvdn_refine(the tail).refine_steps(default 4, 1–49) - used by the refine stage.first_pass_latent(optional) - but required in practice forvdn_refine: the refine path validates that its returned schedule is the real tail of the complete stage's native flow grid, and refuses if it isn't. That provenance check is the thing that stops you from accidentally running a tail that doesn't belong to the pass you think it does.
Outputs: model, sampler, sigmas, stage_context, report_json.
What you keep outside
The description is specific: "Keep learned upscale/reconcile external." VDN stage preparation doesn't do the learned 3D latent upscale or the joint-audio reconcile for you - those stay their own nodes, which is what makes a two-stage VDN graph possible: VDN complete, upscale, reconcile, then a native high stage with an independent model. The pack's 41-vdn-two-pass/ example set has both Full_Stages (run and save both) and Cold_HIGH (load the saved LOW, run only the high stage) variants.
There's also a boundary the author states plainly: VDN temporal Relay needs the separate MiniMaxH3VDNRelayApplyEXPT8 adapter, and low-VRAM side-model persistent identity plus full pretrained/GPU quality are still pending. Read that as "this is a research node with a documented honest gap", not "broken".
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Restart ComfyUI fully and refresh the page. Manager search: MiniMax H3 Audio T8; if the registry copy trails GitHub, install from GitHub - they release independently.
The VDN models are the part that isn't a normal check download. The complete bundle is t8star/Vdn-Minimax-H3-Comfy on HuggingFace, already arranged as your models tree:
hf auth login
hf download t8star/Vdn-Minimax-H3-Comfy --local-dir ComfyUI/models
Access is gated, so request it first. Upstream is OpenVDN/vdn-minimax-h3 if you prefer to assemble it by hand - pin the revision rather than tracking main, since the pack checks shapes and adapter signatures.
Gotchas
Don't stack LoRAs or attention takeovers on a VDN branch. The docs say it directly, for good reason: OpenVDN owns its branch and adapters, and the Composer picks the right Turbo adapter from the base's AdaLN signature. External patches are preserved with warnings rather than rejected outright, but simultaneous speedups aren't promised.
If you see an AdaLN mismatch, the fix isn't renaming files to get past the check. Confirm the bundle has the turbo_pruned_curve_fl2va adapter, and if the error persists, your pruned base has an unsupported curve signature - use the bundle's FL2VA pruned INT8 base, or fall back to the full minimax_h3_fl2va_int8_convrot.safetensors.
Noisy or oddly quiet audio on a VDN run is almost always a step-count, shift or LoRA-identity mismatch. EMA, Ref2VA and OpenVDN turbo LoRAs are not interchangeable. And the licence note matters here more than anywhere: the H3 Community License excludes the US, EU, UK and South Korea from its Applicable Territory - the hosted API is separate, but the local weights are what this node needs.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| av_latent | LATENT | — | |
| stage | COMBO | vdn_complete | 2 options: vdn_complete, vdn_refine |
| refine_steps | INT | 41–49 | — |
| first_pass_latentopt | LATENT | — |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| stage_context | T8_STAGE_CONTEXT | — |
| report_json | STRING | — |