H3→LTX Learned · Bind Separate Refiner (T8 EXP)
An adapter boundary, not an upscaler — the H3→LTX learned handoff
- h3_latent
- original_h3_av
- ltx_video_latent
- model
- noise
- guider
- sampler
- sigmas
- noise
- guider
- sampler
- sigmas
- ltx_video_latent
- stage_boundary
- report_json
What it is
H3 and LTX don't share a latent space, so if you want an LTX refinement pass on top of an H3 take you need a translator. This pack's learned route uses a trained video-latent adapter - no RGB detour, no VAE round trip, no upscaler pretending to be an adapter.
This node is the contract that wraps it. You hand it the adapter's output, the original H3 joint AV, and the external LTX sampler controls you intend to use (MODEL, NOISE, GUIDER, SAMPLER, SIGMAS), plus the two report strings from the adapter and the stage setup. It validates all of it, stamps a stage_boundary with a SHA, and passes the controls straight through.
Nothing is sampled here. It's a signing step. Its outputs are the same noise, guider, sampler, sigmas and ltx_video_latent you fed in, plus stage_boundary and report_json.
Why bother? Because without it, the audit and completion nodes have nothing to check against, and you have no record of which adapter revision produced the render you're looking at three days later.
The video-only part matters
Read the adapter's own contract, not the vibes:
- It converts video latents only. The report declares zero additional upscaler calls and normalises as
normalized_ltx_video. - It states outright that H3 noise masks and reference metadata are not transferred to LTX.
- Audio is not converted. Your H3 joint-AV audio rides through untouched and gets decoded later by the H3 audio VAE, in a separate node. If you were hoping LTX would regenerate the soundtrack, it won't - and H3's native stereo audio is precisely the thing worth keeping.
If you came here looking for the RGB route, where rendered frames are prepared and re-encoded into LTX latent space, that's the sibling pair (...LTXRGBStageBindEXPT8 / ...LTXRGBStageAuditEXPT8). Different reports, not interchangeable.
Inputs and outputs
Configurable-ish inputs: none, really. Everything is a claim being bound - h3_latent, original_h3_av (must be the same object the adapter returned), ltx_video_latent (a LATENT containing only video samples), adapter_report_json, model, noise, guider, sampler, sigmas, setup_report_json.
The one to appreciate is setup_report_json: it's the setup node's report for the LTX stage you're about to run, and it's re-checked at audit time. Change the sigmas or sampler after binding and the audit will refuse the candidate. That's a feature - it turns "why does this look different from yesterday" into an error message.
Outputs: noise, guider, sampler, sigmas, ltx_video_latent, stage_boundary, report_json. Wire the controls and latent into your sampler path, and the boundary forward to the audit.
Typical wiring
H3 AV latent → learned adapter → Bind → (optional LTX EAV / Relay) → SamplerCustomAdvanced or Stage Sample
└─ stage_boundary ─────────────────────────→ Audit
original_h3_av ──────────────────────────────────────────────────────→ Original Audio Decode
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Manager: search MiniMax H3 Audio T8, install, then fully exit ComfyUI and restart and refresh the browser. Node packs this large don't hot-load; the pack's own guide says if everything goes red or the nodes look missing, update ComfyUI core, the frontend and Manager together rather than just re-pulling this repo.
H3 side: transformer in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, video and audio VAEs in models/vae. The adapter, LTX-2.x weights and its encoder are separate downloads following their own licences. requirements.txt deliberately pulls in nothing, so installation can't overwrite your torch/CUDA stack. The pack contains no weights at all.
And the standing caveat for this whole family: H3's local weights are licensed only outside the US, EU, UK and Korea, outputs included.
Common issues
"Learned LTX output must contain only video samples." You passed an AV latent, or a latent with extra keys. The adapter returns video; the audio lives elsewhere.
"Masks are not transferred" surprises you at render time. H3 reference/mask metadata doesn't survive into the LTX stage. If your H3 pass used mask-based references, expect the LTX refine to be a straight video refinement of what came out, not an extension of that mechanism.
Different render, same settings. Something upstream changed - adapter revision, LTX weights, LoRA stack. The audit's SHA comparison is how you prove it.
Red nodes after an update. Update core + frontend + Manager, restart fully. This route also expects you to start from the pack's bundled workflow rather than hand-rolling the graph.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| h3_latent | LATENT | — | |
| original_h3_av | LATENT | — | |
| ltx_video_latent | LATENT | — | |
| adapter_report_json | STRING | — | |
| model | MODEL | — | |
| noise | NOISE | — | |
| guider | GUIDER | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| setup_report_json | STRING | — |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| noise | NOISE | — |
| guider | GUIDER | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| ltx_video_latent | LATENT | — |
| stage_boundary | T8_LTX_LEARNED_STAGE_BOUNDARY | — |
| report_json | STRING | — |