Nodes/MiniMax H3 Activation Chunk - Star7/MiniMax H3 分块解码 - Star7
ComfyUI Node

MiniMax H3 分块解码 - Star7

Your H3 latent is two latents wearing a trenchcoat — this node gets both back out

By star7code·Created 2 months ago·Updated 5 days ago· 86
MiniMax H3 分块解码 - Star7
  • av_latent
  • video_vae
  • audio_vae
  • frames
  • audio
  • report

MiniMax H3 doesn't generate a silent video and bolt audio on afterwards - it generates both streams together, so its sampler output isn't one latent. It's a nested tensor holding a video latent and an audio latent side by side. A plain VAEDecode node expects one .samples tensor and one VAE, which is exactly the wrong shape for that. MiniMaxH3ChunkedDecodeStar7 is the small, boring node that unpacks the pair and sends each half to the decoder that actually understands it.

You'd reach for it at the very end of an H3 chain, after the sampler or after the one-click HD node, instead of wiring two decoders by hand.

What it actually does

Give it the audio-video latent, the H3 video VAE and the H3 audio VAE, and it does four things:

  • unbinds the nested latent and refuses anything that isn't one - a plain image latent gets a clear "expected a nested video/audio latent" error rather than a shape crash three nodes later;
  • decodes the video stream through the video VAE and the audio stream through vae_decode_audio, so you get real IMAGE frames and a real AUDIO object;
  • flattens any 5-D [B,C,T,H,W] decoder output into a normal image batch and checks the frames are finite, in 8-frame slices so it never allocates a full UHD-sized boolean tensor just to test them;
  • hands you a short report string saying what it did.

About that name. "分块解码" is not this node inventing its own chunking scheme. The chunking happens inside the H3 VAE itself, which decodes with temporal streaming and spatial tiling so a long clip doesn't materialise every frame's activations at once. What the node adds is detection: it looks at the connected video VAE, recognises the native MiniMaxH3VideoVAE, reads its tiling flag and tile size, and tells you which route you got - native H3 temporal streaming + spatial tiles (Npx) or the fallback VAE-managed decode. Worth reading that line, because if you've swapped in a non-native VAE the report is the only place that says so out loud.

Its tiling is also completely unrelated to the tiling checkbox on the HD node, which tiles sampling predictions. One is decode memory, the other is refine memory.

Inputs and outputs

Three inputs, no hidden knobs:

  • av_latent - the nested H3 latent from your sampler (or from MiniMaxH3OneClickHDStar7).
  • video_vae - the H3 video VAE.
  • audio_vae - the H3 audio VAE. Yes, it's a separate VAE, and yes, forgetting it is the usual mistake.

Outputs are frames (IMAGE), audio (AUDIO) and report (STRING). Wire frames and audio into your video combine node for a clip with sound; report goes to a text display if you want it, otherwise ignore it. The node isn't an output node, so nothing is saved until you connect a save/combine node yourself.

Install

It ships inside the Star7 MiniMax H3 Activation Chunk pack, so you install the pack once:

cd ComfyUI/custom_nodes
git clone https://github.com/star7code/minimax-h3-chunk-star7.git

Or search MiniMax H3 Activation Chunk - Star7 in ComfyUI Manager, or comfy node install minimax-h3-chunk-star7. Restart afterwards. The pack's pyproject.toml pulls scipy, scenedetect and ultralytics - those serve the pack's other nodes (face repair, DLSS enhancement), not this one, but they land in your environment regardless. This node needs no model download of its own: it uses the H3 VAEs you already have. The rest of the pack is a deep rabbit hole of attention backends (Comfy Kitchen INT8, SLA, Sol, step-level Hybrid) and QKV/RoPE/MLP activation chunking aimed at running H3 at long durations on 20-series-and-up cards; the author's own numbers are from a 2080 Ti 22GB.

Worth one line on context: H3 is a 33B omni-modal model whose community licence excludes the US, EU, UK and South Korea from local weights. That's a licensing problem, not a technical one, but it's real.

Where it bites

The system RAM bill, not VRAM. The normal IMAGE output has to hold every decoded frame as RGB float32 - roughly 12 bytes per pixel per frame. A 10-second 1MP clip is about 2.8 GB; a 4K clip of the same length is closer to 24 GB. The node estimates this and puts it in the report as final IMAGE tensor≈N MiB RAM, which is the most useful thing in that string. If your machine starts swapping at decode time, this is why, and no amount of activation chunking upstream will help.

Wrong latent. Feed it a plain image latent and it errors on purpose - nested H3 audio-video only.

Corrupt output. If the VAE returns NaN or Inf frames it raises instead of pushing garbage into your muxer. Annoying when you're tired, correct when you're not.

Reading the log. Everything it logs is one line, prefixed Star7 H3 decode. The route and the RAM estimate are both there. If something's off, that line is what to paste into an issue.

CategoryStar7/MiniMax H3

Inputs (3)

NameTypeDefaultDescription
av_latentLATENT—
video_vaeVAE—
audio_vaeVAE—

Outputs (3)

NameTypeDescription
framesIMAGE—
audioAUDIO—
reportSTRING—