H3→LTX Learned · Audit Refiner Candidate (T8 EXP)
The audit that stops you lying to yourself about the LTX refiner
- stage_boundary
- h3_latent
- original_h3_av
- ltx_video_latent
- model
- noise
- guider
- sampler
- sigmas
- candidate_latent
- candidate_latent
- original_h3_av
- duration_seconds
- report_json
Why you'd want this
H3 gives you a joint video+audio pass. LTX-2.x gives you speed and a different look, and the community has spent a year treating it as the second stage in local video pipelines - the whole "draft in one, refine in the other" habit. This pack's H3→LTX learned route is the version of that idea where H3's video latent is converted to LTX latent space by a trained adapter rather than by decoding to RGB and re-encoding.
It works. It's also very easy to mis-wire, because the graph hands you six or seven strings and objects that all have to agree, and the failure mode of a mis-wired refiner is a plausible-looking render, not a crash.
This node is the end-of-chain check. It takes the boundary MiniMaxH3LTXLearnedStageBindEXPT8 gave you, every value that went into it, and the candidate_latent that came out of your external LTX sampler - then recomputes the contract and compares. Mismatch, and it raises.
What it actually verifies
- That
original_h3_avis the same object ash3_latent. The adapter is contractually supposed to hand back the unchanged source H3 AV; if it quietly produced a new one, this catches it. - That the adapter report is the owned, video-only conversion: the right schema, the expected model SHA and revisions, output normalisation of
normalized_ltx_video, zero additional upscaler calls, and the stated mask policy that H3 noise masks and reference metadata are not transferred to LTX. - That the learned LTX latent has the geometry the adapter claims - 128 channels, the timeline's latent frame count, and height/width derived from the H3 video latent.
- That the connected MODEL, NOISE, GUIDER, SAMPLER, SIGMAS and setup report are byte-for-byte the ones bound earlier, and the candidate latent's shape matches the adapter output.
And the honest bit, which is why I like this node: the report says sampler_execution_proven: false. It checks what you connected and what shape came out. It does not prove your external LTX sampler actually ran. For that you want the Certified variant plus the Sample node's live proof.
Inputs and outputs
You don't configure anything here - every input is one of the claims being re-checked: stage_boundary, h3_latent, original_h3_av, ltx_video_latent, adapter_report_json, model, noise, guider, sampler, sigmas, setup_report_json, candidate_latent. Pass-through everywhere.
Outputs are candidate_latent (to your LTX video decode - TAEHV or whatever your workflow names), original_h3_av (to MiniMaxH3LTXOriginalAudioDecodeEXPT8, because the audio never went through LTX at all), duration_seconds and report_json. Read the report; it's the only place audio_policy and the boundary text are stated.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Or search MiniMax H3 Audio T8 in ComfyUI Manager, then fully exit and restart ComfyUI and refresh the page. Manager publishing lags GitHub, so if the node is missing while GitHub has it, clone.
H3 models: transformer in models/diffusion_models, Qwen3-VL text encoder in models/text_encoders, video and audio VAEs in models/vae. LTX-2.x weights, its text encoder and the learned adapter are separate downloads with their own licences. The pack installs no Python packages of its own - requirements.txt exists only to say so.
Start from a working learned-route graph rather than wiring this by hand; the pack keeps those under examples/workflows/35-h3-ltx-latent. And note H3's local weights are geofenced out of the US/EU/UK/Korea by the MiniMax H3 Community License, outputs included.
Common issues
"Learned adapter report is not the owned video-only conversion." You fed it the RGB prep report, or a stale adapter report from a previous run. The learned and RGB routes have different reports and don't interchange.
"Requires an explicit audio reconciliation stage." The adapter changed the timeline length, so the original H3 audio no longer lines up. The pack refuses to let you silently mux a shorter/longer track; fix the duration deliberately.
"Source, adapter report or sampler controls changed." You edited something upstream - new seed, different sigmas, swapped sampler - after binding. That's the intended behaviour; the audit is telling you the graph is no longer the thing you signed.
Red nodes. Update ComfyUI core, frontend and Manager together, then restart fully.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| stage_boundary | T8_LTX_LEARNED_STAGE_BOUNDARY | — | |
| h3_latent | LATENT | — | |
| original_h3_av | LATENT | — | |
| ltx_video_latent | LATENT | — | |
| adapter_report_json | STRING | — | |
| model | MODEL | — | |
| noise | NOISE | — | |
| guider | GUIDER | — | |
| sampler | SAMPLER | — | |
| sigmas | SIGMAS | — | |
| setup_report_json | STRING | — | |
| candidate_latent | LATENT | — |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| candidate_latent | LATENT | — |
| original_h3_av | LATENT | — |
| duration_seconds | FLOAT | — |
| report_json | STRING | — |