H3→LTX Learned · Decode Original H3 Audio (T8 EXP)
Your H3 audio survives the LTX pass — here's how to get it back
- original_h3_av
- audio_vae
- original_h3_audio
- report_json
What it is
The H3→LTX refinement pass is video-only. LTX never sees your H3 audio, the adapter never converts it, and there's no LTX audio latent to decode at the end of that chain. What you have instead is the original H3 joint-AV latent, passed through unchanged, sitting there next to your refined video.
This node turns it into sound. It decodes only that original H3 audio latent with the H3 Audio VAE - the same VAE the pack uses everywhere else for H3 audio - and hands you a normal AUDIO object you can feed to CreateVideo, a save-audio node, or your muxer.
The name is literal, and so is the discipline: never reinterpret an H3 audio latent as LTX audio, and never decode the H3 video latent here. If you hand it a video-only H3 input it fails explicitly rather than guessing, which sounds annoying until you remember the alternative is a silent, wrong-sounding track.
Why this matters more than it looks
H3's selling point is that audio is generated jointly with the video - not bolted on afterwards. That's genuinely the good part; the community's reaction on H3 release was basically "finally, a local model with usable native audio." Throwing that away to get an LTX look would be a strange trade. Keeping the original track while refining the picture is the sane pipeline, and this node is the piece that makes the keep explicit rather than accidental.
It also enforces the timing rule that people get wrong: if the adapter changed the timeline length (audio_duration_adjustment_required in its report), this node refuses. Trim/pad needs a deliberate audio reconciliation step, not a hopeful mux where a 4.7-second track gets glued to a 5-second clip and everyone pretends not to notice.
Inputs and outputs
Three inputs, and the first one is the whole trick:
original_h3_av- the H3 joint-AV latent. In a properly wired learned-route graph this is the same object the Bind node passed through; you don't decode anything to produce it.adapter_report_json- the adapter's report string. It's used to re-derive the expected H3 frame count, temporal policy and reference-prefix latents, so a stale report is a real failure mode.audio_vae- a VAE input. Load the H3 audio VAE, not the video one. Feeding the video VAE is the kind of mistake that produces noise and a long evening.
Outputs: original_h3_audio (AUDIO) and report_json. The report states the policy flatly - H3_audio_VAE_only_no_LTX_audio_conversion - which is exactly what you want on record when a colleague asks where the sound came from.
Where it fits
H3 joint AV → learned adapter → Bind … Sample … Audit
↘ original_h3_av → Original Audio Decode → AUDIO
refined ltx_video_latent → TAEHV decode → images ────────────────→ CreateVideo(audio=…)
Notice you can wire it from either the audit's original_h3_av output or straight from the source latent - both are the untouched original.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Manager users: search MiniMax H3 Audio T8, install, then fully quit and restart ComfyUI and refresh the page. The pack's guide notes that if nodes turn red or vanish, the fix is usually updating ComfyUI core, the frontend and Manager together - pulling only this repo often isn't enough, because these nodes follow native H3 APIs in core.
Required files: H3 diffusion model in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, and both the video and audio VAEs in models/vae. LTX weights and the adapter are separate downloads. The repo itself never ships weights, and it installs no Python packages - requirements.txt is deliberately package-free to protect your torch/CUDA install.
Common issues
Output is hiss, silence, or a low rumble that isn't speech. You almost certainly loaded the H3 video VAE into audio_vae, or passed a video-only latent. Both produce an unhelpful result rather than a clean error in some graphs.
"Audio duration adjustment required." The LTX side changed the length of the clip. Decide deliberately: trim the video back, or pad/extend the audio in an explicit node. The pack won't hand you a mismatched pair.
Track sounds fine but drifts late in the clip. Check for a duration change again. A tiny mismatch at the start is easy to miss and obvious at the end.
Nothing renders / nodes red. Update core + frontend + Manager and restart fully. And remember the H3 licence question - the local weights are only licensed outside the US, EU, UK and Korea, outputs included.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| original_h3_av | LATENT | — | |
| adapter_report_json | STRING | — | |
| audio_vae | VAE | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| original_h3_audio | AUDIO | — |
| report_json | STRING | — |