H3 Decode Video
Decoding a 60-second MiniMax H3 run without eating your RAM
- latent
- vae
- audio_vae
- video
- frames
- report
MiniMax H3 opened its weights in August 2026 - a 33B omni-modal video model, 4–15 second clips at up to 2K/24fps with native stereo audio, which is the thing open video was missing. The catch, if you're in the US, EU, UK or South Korea: the community licence excludes your territory entirely.
You get 4–15 seconds, so to get a minute you run the pack's H3 extend loop - sample a segment, append it, carry on - and the loop hands you one big latent with every scene in it. H3 Decode Video is what comes after. It turns that latent into frames on disk, one scene at a time, so a 60-second multi-scene run never exists in RAM as a single batch.
Why you'd reach for it
The naive path is VAE Decode, and it works right up until it doesn't. Decoding six or eight appended H3 scenes in one call means one enormous image batch in memory, and on a 16GB card that's the crash you get instead of a file. Decoding each segment inside the loop instead means stitching everything yourself, and the joins would show - a segment's first frames are generated against the frames before it and carry a little of them.
H3 Decode Video does it correctly and cheaply: one node, one decode pass, frames landing on disk as they come, and a clean first frame at every cut.
How it works
The latent from H3 Extend Append isn't only video - it carries the audio the model generated alongside it. The node splits the two, then works out where the scenes are: the loop records which token spans were sampled as fresh scenes, and where a segment's ends reveal that a cut trimmed it, that becomes a span of its own.
Each scene is then decoded on its own. A long one is decoded in pieces of up to twelve clips, each piece given a lead-in clip of context that's thrown away afterwards, so pieces join without a visible seam and the first frames of a scene that opens on a cut carry nothing from the scene before. Frames are appended to a frame cache on disk as they're produced - frames.bin plus a small cache.wasframes manifest - and the audio is decoded once, through audio_vae, into the same cache.
Inputs and outputs
You set three things, really.
- latent - the finished clip from the extend loop, i.e. H3 Extend Append's output.
- vae - the H3 video VAE. Wire the wrong VAE and you'll be told the latent isn't an H3 video+audio latent, which is the node's way of saying "this isn't an H3 clip".
- root -
temp(cleared when ComfyUI restarts) oroutput(kept, and readable by Load Video Cache later). If you want the clip tomorrow, pickoutput. - name - the cache's name below root, numbered per run:
frame_cache/episodebecomesframe_cache/episode_00001. - delete_after_save and the optional audio_vae - sound only exists if you feed the audio VAE.
Out comes video, a clip reader that Save Video (Advanced) consumes a frame at a time; frames, an INT with the count decoded; and report, a STRING that lists each scene's frame count and where it opens in the clip. Don't ignore the report - if you're cutting scenes apart later, it's the map, and it's the only place that information is written down.
Install
ComfyUI Manager → search WAS Node Suite v3. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+, Python 3.10+, restart, done. There's no model download and no pip step for the suite itself - v3 installs zero packages, unlike v2, which was famous for its dependency fights and "import failed" posts after every ComfyUI update. The pack ships example H3 graphs under docs/workflows/ - start from minimax-h3-extend-loop.json.
Common issues
delete_after_save: true re-decodes on every queue. Deliberately - the node's fingerprint is invalidated so it can rebuild the cache and hand a clean copy to the save, then delete it. That's the one-shot-mode setting. Set it false while you're still iterating, or you'll pay for the decode every single time you press Queue.
It's a disk trade, not a free lunch. The cache is raw frames: 8-bit codes by default, half floats when the source needs it. A minute of 24fps 2K footage is real gigabytes, and every run makes a new numbered folder. Clean up output/frame_cache/ occasionally.
A cache in temp doesn't survive a restart, so Load Video Cache can't rescue it. The node does tell you where the cache went in the log line, which is the fastest way to find it when you need to check.
These are new nodes. They landed in the pack's most recent commit, so there's essentially no community write-up, no tutorial, no StackOverflow trail yet. If something looks wrong, the report output and the ComfyUI console are your diagnostics.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| latent | LATENT | The finished clip, from the loop H3 Extend Append builds. | |
| vae | VAE | The H3 video VAE. | |
| root | COMBO | temp | Which folder the cache lands in: 'temp' = cleared when ComfyUI restarts; 'output' = kept until deleted, for Load Video Cache. |
| name | STRING | frame_cache/h3 | The cache's name below root, numbered on each run, as `frame_cache/episode` for `frame_cache/episode_00001`. |
| delete_after_save | BOOLEAN | false | `true` = the cached frames are deleted once a save has written all of them, and the decode runs again on every queue; `false` = they are kept. |
| audio_vaeopt | VAE | The H3 audio VAE, for the clip's sound. Left empty, the clip is silent. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | The decoded clip at 24 fps, read from disk by Save Video a frame at a time. |
| frames | INT | Frames decoded. |
| report | STRING | Each scene's frames and where it opens in the clip. |