Nodes/ComfyUI_JR_MiniMaxH3Node/JR MiniMax H3 Split AV Latent
ComfyUI Node

JR MiniMax H3 Split AV Latent

Splitting an H3 AV latent so you can actually edit the video stream

By Goldlionren·Created 21 days ago·Updated about 9 hours ago· 28
JR MiniMax H3 Split AV Latent
  • av_latent
  • video_latent
  • audio_latent

Here's the thing about the packed H3 AV latent: it's a NestedTensor with a video stream and an audio stream bundled together, and most standard latent nodes have no idea what to do with it. You can't run it through a spatial upscaler, you can't interpolate it, you can't inspect one stream without the other. JR_H3_SplitAVLatent is the un-bundler: it takes an official H3 AV LATENT and splits it into a standalone video_latent and audio_latent, both as ordinary ComfyUI LATENT dictionaries.

The design is deliberately cheap. The split uses the public unbind() on the nested tensor, outputs reference the original tensors directly - no cloning, no casting, no device moves, no forced contiguous(). That's not laziness, it's correctness: copying a 33B-class model's latent stream just to split it would be a waste of VRAM for zero benefit. The validation is where it spends its effort - it checks stream count, tensor types, video [B,24,T,H,W], audio [B,32,2,T], matching batch and finite values.

What it unlocks

The whole point is the two streams diverge from here:

  • video_latent → spatial latent ops (upscaling, interpolation, resizing) → back into an AV builder.
  • audio_latent → carried through untouched → back into an AV builder.

The canonical pattern is the two-pass upscale: split the sampled AV latent, upscale only the video stream with the Neural Latent Upscaler, then rebuild with the AV Latent Builder and run a pass-2 sampler. You can also Save Latent each stream independently, Load Latent them in another workflow, and rebuild - that's the persistence path for partially-processed generations.

The one rule that burns people

The audio latent has no spatial axes. It's [B,32,2,T] - no H, no W. Spatial upscalers and interpolators can only ever touch the video_latent output. Feed the audio stream into a spatial chain and you'll get a shape error or silently meaningless output. Keep them separate until you rebuild.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/Goldlionren/ComfyUI_JR_MiniMaxH3Node
<your-comfyui-python> -m pip install -r ComfyUI_JR_MiniMaxH3Node/requirements.txt

or ComfyUI Manager → search ComfyUI_JR_MiniMaxH3Node. Pack-wide install; same Python as ComfyUI, restart.

Troubleshooting

  • "Not a NestedTensor" or stream-count errors - the input must be an official H3 AV latent from the native conditioning/sampler path. A plain image or video latent won't split.
  • Batch mismatch between streams - the builder paired them, so if you've been moving things around between split and rebuild, keep the batch axes consistent.
  • Video stream looks wrong after spatial ops - confirm you only ran spatial nodes on video_latent, and that any save/load round-trip preserved the [B,24,T,H,W] layout. Contiguity is handled by the save node's own path, so that's not your problem to babysit.
CategoryJR MiniMax H3/Latent

Inputs (1)

NameTypeDefaultDescription
av_latentLATENT

Outputs (2)

NameTypeDescription
video_latentLATENT
audio_latentLATENT