H3 Low VRAM
One node, long clips on the card you already own
- model
- model
MiniMax H3 is a 33B omni-modal video model with native audio, and the honest reason people bounce off it locally is memory, not quality. Long clips mean very long token sequences, and the sequence length is what eats VRAM, not the weights alone. H3 Low VRAM is the smallest useful node in the pack for that: one MODEL in, one MODEL out, no settings. Drop it after your H3 loader and any LoRA, and every H3 sampler downstream runs the same way - all the transformer blocks except attention processed in slices of 8192 tokens, with only the weights that fit beside the clip left resident and the rest streamed in as each block runs.
The author's claim is "the same output at a lower memory peak", and that is the right way to read it. This is not a quality/speed tradeoff knob like a quantised checkpoint. It is a scheduling change.
Why it works at all
Attention has to see the whole sequence, so slicing it is a no-op - which is exactly why the node leaves attention alone and slices everything else. A block's feed-forward, norm and projection work is per-token, so it can be broken into chunks with no numeric difference; you just need the weights in memory when a chunk needs them and you can hand them back afterwards.
The pack works its budget out the tidy way: it measures free VRAM inside each model call, sets aside roughly 192 KB per token plus 3 GB of fixed working memory, and keeps a further 768 MB free as margin. So on a busier graph the budget tightens rather than the run falling over.
Inputs and outputs
There is one input, model - the MiniMax H3 model after any LoRA. There is one output, model, for every H3 sampler in the graph.
That triviality is the point, and also the trap. Because there is nothing to configure, people wire it in the middle of a chain and wonder why nothing changed. Place it as the last thing you do to the model before sampling:
Load H3 model → LoRA Loader(s) → H3 Low VRAM → [H3 Tiled Sampler / H3 Tiles / KSampler]
Put a LoRA after it and the LoRA loader is working on a patched clone, which is where you get confusing behaviour. Put it before the model loaders and it has nothing to wrap.
Note the same mechanism is available inside other nodes - H3 De-RoPE Stretch exposes a low_vram boolean for its own pass, and H3 Tiles applies the tiling sibling of this idea to every model call. If a whole graph is H3, one H3 Low VRAM node early is cleaner than three of them scattered around.
Install
ComfyUI Manager, search WAS Node Suite v3, install, restart. Or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
ComfyUI 0.14.0+ and Python 3.10+. Nothing is pip-installed by the pack - requirements.txt carries a comment and nothing else, and the pack never runs pip on its own. No model downloads for this node; it wraps whatever H3 checkpoint you already loaded.
Two things worth knowing from the pack's own notes: the first start after installing (and after every update, since the bytecode is recompiled) takes an extra moment, and config.yaml - written under your ComfyUI user directory in was-node-suite/ - has features.network: false by default, so nothing reaches out to the internet on your behalf.
When it does not help
If your OOM is at VAE decode or during the audio pass rather than in the sampler, this node will not save you - it only changes how transformer blocks are executed. That is what the VAE-side nodes in the H3 family are for.
If you are loading a permissively-licensed video model instead of H3, none of this applies: Wan 2.2 and LTX have their own memory tooling in ComfyUI core, and H3 Low VRAM is written for the H3 patcher specifically. And if you are in the US, EU, UK or Korea, the licence, not the memory, is your blocker - H3's community licence excludes those territories from running the local weights at all.
One more practical note: expect the run to be slower, and expect it to be identical otherwise. If your output changed after adding this node, something else in your graph changed too.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | The MiniMax H3 model, after any LoRA. | |
| head_chunks | INT | 00–56 | 0 = off; 4 = attention runs over 14 of the 56 heads at a time; 8 = 7 at a time. Every head is computed as before, so the result is unchanged. Frees roughly a third of the memory attention holds on a long clip. |
| query_chunks | INT | 00–64 | 0 = off; 4 = attention answers the clip's tokens in 4 chunks against keys and values worked out once. Each token is computed as before, so the result is unchanged. Frees about half the memory attention holds, at the cost of one extra attention projection per block. Combines with head_chunks. |
| token_slice | INT | 81921024–65536 | 8192 = norms, projections and feed-forward run 8192 tokens at a time; 4096 holds about half the working memory there. The result matches up to float rounding. Lower it when memory is still short after head_chunks and query_chunks. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| model | MODEL | The model for every H3 sampler in the graph. |