H3 → LTX 视频潜空间 / Learned Adapter (T8 EXP)
Converting a latent without touching a single pixel
- h3_latent
- ltx_video_latent
- original_h3_av
- output_frames
- output_fps
- report_json
MiniMax H3 is the heavyweight: 33B, joint picture-and-audio, great motion, slow per second of video. LTX is the opposite trade - Lightricks' speed-first family, where 2.3 rebuilt the VAE and shipped a new latent space aimed at sharper texture and edge detail. The obvious pairing is "generate in H3, refine in LTX" - and they don't share a latent space. Normally that means decoding to RGB and re-encoding into LTX: a lossy round trip through pixels for no good reason.
MiniMaxH3LTXLatentAdapterEXPT8 is the no-round-trip path. It takes a standard Comfy LATENT from H3, runs it through a learned adapter that maps H3's video latent into LTX's latent grid, and hands you back LTX-shaped video latent. No RGB detour, no second upscaler slapped on the end. It's marked EXP for a reason you'll notice immediately: the frame maths and the assets are pinned hard.
What it does with your tensor
H3 packs video as B,24,T,H/16,W/16, and when you're using a joint AV route the latent is a nested object with the audio branch as B,32,2,T alongside it. The node splits that, converts only the video branch, and returns the original H3 object untouched on the second output (original_h3_av). It explicitly refuses to reinterpret H3's audio as LTX audio, and it doesn't touch masks or reference conditioning. If you were hoping for a one-node "H3 clip becomes an LTX project", this is a third of that.
The adapter itself is a ~195M-parameter learned model - real weights, not a resize. And those weights are authenticated, not just named: the source hashes of five files in the upstream h3_ltx_adapter module, plus the SHA-256 of config.json and model.safetensors, are baked into the node. Wrong revision, refuses to run. Right revision, runs from the bytes it checked.
The temporal contract is the whole story
H3 emits frames on a 17n+5 grid; LTX wants 8n+1. Those grids only coincide at some counts, and 73 frames is the happy one:
source_frames = 73,frame_policy = exact→ 73 frames in, 73 out, audio duration unchanged.source_frames = 124,exact→ explicit rejection. Not a silent crop, not a guess. Pickpad_to_ltx_grid(129 frames, and you now have audio to extend or trim yourself) orcrop_to_ltx_grid(121 frames, and you trim the audio).- Anything else - rejected with "Expected an explicit native H3 17n+5 frame count". The node will not infer duration from the latent shape.
source_fps must be 24. Not "roughly 24", not "we'll resample" - the recipe is CFR24 and the node says so. And the incoming canvas has to already be 32-pixel aligned, because the adapter doesn't stretch.
device and precision default to cpu / float32 - conservative and slow. CUDA/BF16 is an explicit opt-in, not an automatic speedup, and the adapter fitting in ~460MB of VRAM says nothing about whether your whole LTX refiner route fits. There's also reference_prefix_latents for H3 latents that carry reference frames as a prefix - tell it exactly how many latent frames are prefix, because it's an offset, not a filter. normalization stays on comfy_normalized unless you specifically generated a raw-H3 latent.
Outputs: ltx_video_latent (feed your LTX path), original_h3_av (your H3 audio and picture, unmodified), output_frames and output_fps as numbers, and a report_json string carrying the asset revisions and adapter precision.
Install and assets
The node comes with the pack - ComfyUI Manager, search MiniMax H3 Audio T8, or:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Then fully restart. The pack's requirements.txt is deliberately empty, so nothing here will stomp your Torch/CUDA install.
The two external pieces you must supply yourself - no download, no bundling - are the Sana source checkout at the pinned revision, specifically models/minimax_h3/Sol-H3-Spark/runtime/stage2_ops/h3_ltx_adapter, and the Efficient-Large-Model/H3-to-LTX-Latent-Adapter model at its pinned revision with the original config.json and model.safetensors. Type both paths into source_directory and model_directory. The template in examples/workflows/35-h3-ltx-latent converts and saves - it is not a working refiner; the downstream LTX conditioning, text cache and refinement pass are yours to prepare.
Where it'll stop you
Expect "Missing or oversized adapter source" or a hash mismatch if you cloned Sana at main instead of the pinned commit - the most common way to fail this node before it does anything interesting.
Expect "Source frame count and exact reference prefix do not match H3 temporal shape" if source_frames isn't the actual H3 output count. Count it, don't assume the latent encodes it.
And budget your audio handling. The node doesn't mux, doesn't use -shortest, and doesn't touch the soundtrack - after a pad or crop your video and audio lengths legitimately disagree, and it's on you to reconcile them before delivery. The pack's own answer is to handle the tail explicitly rather than let a muxer clip it.
Inputs (10)
| Name | Type | Default | Description |
|---|---|---|---|
| h3_latent | LATENT | — | |
| source_frames | INT | 73 | 实际H3输出帧数17n+5,不从潜空间猜时长。 |
| source_fps | FLOAT | 24.00 | 当前配方仅24fps;不自动重采样。 |
| frame_policy | COMBO | exact | exact保持帧数;124帧需显式选pad→129或crop→121,原音需另行对应处理。 |
| source_directory | STRING | 固定Sana源码中h3_ltx_adapter目录,非整个仓库。 | |
| model_directory | STRING | 包含官方config.json和model.safetensors的目录。 | |
| device | COMBO | cpu | 2 options: cpu, cuda |
| precision | COMBO | float32 | 2 options: float32, bfloat16 |
| reference_prefix_latents | INT | 0 | — |
| normalization | COMBO | comfy_normalized | 2 options: comfy_normalized, raw_h3 |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| ltx_video_latent | LATENT | — |
| original_h3_av | LATENT | — |
| output_frames | INT | — |
| output_fps | FLOAT | — |
| report_json | STRING | — |