H3→LTX RGB · Bind Separate Refiner (T8 EXP)
Frames you already have, audio you keep
- source_frames
- source_audio
- prepared_frames
- ltx_latent
- model
- noise
- guider
- sampler
- sigmas
- noise
- guider
- sampler
- sigmas
- ltx_latent
- stage_boundary
- report_json
What it is
There are two ways into an LTX refinement pass over H3 footage in this pack, and they are not the same thing:
- Learned route - a trained adapter converts the H3 video latent straight into LTX latent space. No pixels involved.
- RGB route - you take rendered frames, prepare them (resize/crop/what the workflow says), encode them into LTX latent space, and refine. This is the classic route, and it works on any footage, including video H3 didn't produce.
This node is the RGB route's binding step. It takes the original H3 source frames, the original H3 AUDIO, the prepared frames, the prep report, the encoded LTX latent, and the external refiner controls (MODEL, NOISE, GUIDER, SAMPLER, SIGMAS) - validates that the two paths agree, and emits a signed stage_boundary.
Notably: it validates the audio as a bypass. The original H3 audio is passed through untouched and never converted, because the LTX refiner only touches video. That explicit bypass is the design; if you were expecting the refiner to regenerate or improve sound, that's not what's happening anywhere in this pair.
No sampling here, and the node is careful to claim no cache either - it binds, it doesn't execute.
Inputs and outputs
Inputs worth understanding:
source_framesandsource_audio- the original H3 result. The frames are the reference geometry the prep must respect; the audio is the thing being preserved.prepared_frames- the output of the RGB preparation step. Shape here must line up with the source per the prep report.prep_report_json- the preparation node's report. This is one of the strings the audit will re-check, so keep the pair together.ltx_latent- the already-encoded, already-lifted LTX latent.model,noise,guider,sampler,sigmas- your external refiner controls, passed straight through.setup_report_json- the LTX stage setup's report.
Outputs: noise, guider, sampler, sigmas, ltx_latent, stage_boundary, report_json. Feed the controls and latent into your sampler; feed the boundary to MiniMaxH3LTXRGBStageAuditEXPT8.
Which route should you pick
Learned if you have the H3 latents and the adapter weights and want the fastest, cleanest handoff. RGB if your source is a finished clip - an older H3 render, someone else's video, anything you didn't just generate in this graph - or if you want to see and preprocess the actual frames before refinement.
Practical note from the community's own history with LTX: it's the fast one, and people accept its quality because of the speed and because it's the local video model that has audio at all. As a refinement stage on top of H3 output you're mostly buying a different look and a different temporal feel, not a magic fidelity lift. Expect to tune sigmas and use a shallow schedule.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
ComfyUI Manager: search MiniMax H3 Audio T8, install, then fully exit and restart ComfyUI and refresh the page. If Manager's copy lags GitHub, clone - the registry and the repo release separately.
Needed on disk: H3 transformer in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, H3 video and audio VAEs in models/vae, plus the LTX-2.x weights, its text encoder and whatever "learned lift" assets the workflow names. No weights ship with the repo, and requirements.txt installs nothing by design so it can't clobber your torch/CUDA build.
Start from examples/workflows/60-ltx-rgb-stage-split (the S27 ... StageBound_Separate_EXP graphs) instead of wiring this yourself. Those graphs exist in plain and Identity-Preserve flavours and the two are not interchangeable.
Common issues
"Learned LTX output must contain only video samples" style complaints aren't the errors you'll see here. On the RGB route the failure is geometry: prepared frames that don't match the source per the prep report, or an ltx_latent that doesn't correspond to the prepared frames. Re-run the prep node rather than nudging values.
You changed the sampler controls after binding and the audit failed. Intended. Rebinding is one extra click and it's how the audit can tell you the candidate belongs to the stage you signed.
Frames look correct but the final file refuses to decode. The pack's own development record is candid about this route: there have been runs where the latents matched and the encoded media still failed strict H.264 decoding. Treat "the graph completed" and "the file is good" as two different facts, and always play the export.
Red nodes. Update ComfyUI core, frontend and Manager together, then restart. These EXP nodes track moving native H3 APIs.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| source_frames | IMAGE | — | |
| source_audio | AUDIO | — | |
| prepared_frames | IMAGE | — | |
| prep_report_json | STRING | — | |
| ltx_latent | LATENT | — | |
| model | MODEL | — | |
| noise | NOISE | — | |
| guider | GUIDER | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| setup_report_json | STRING | — |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| noise | NOISE | — |
| guider | GUIDER | — |
| sampler | SAMPLER | — |
| sigmas | SIGMAS | — |
| ltx_latent | LATENT | — |
| stage_boundary | T8_LTX_RGB_STAGE_BOUNDARY | — |
| report_json | STRING | — |