Extensions/Unload Models
ComfyUI Extension

Unload Models

Passthrough nodes that free models ComfyUI's memory manager will not free, so large models fit on small cards.

By neezoy·Created 2 months ago·Updated 2 months ago· 2
neezoy/ComfyUI-UnloadModels
Nodes—
On cloudLocal install
Stars2
Updated2 months ago
Readme

ComfyUI-UnloadModels

Two passthrough nodes that free models ComfyUI's memory manager leaves loaded.

I wrote these to run the full 21 GB MiniMax H3 video model on a 16 GB card. Before them the machine froze on the first render. After them it renders, and each step got 3x faster.

| configuration | s/it | |---|---| | stock ComfyUI | 33.75 | | with Unload CPU Models | 11.00 |

Measured on a Radeon RX 9070 XT (16 GB, gfx1201), ROCm 7.1, ComfyUI 0.33.0, MiniMax H3 ref2va at 0.4 megapixels (608x352), 5 s, 20 steps, res_multistep. Same seed and resolution in both runs.

The problem

ComfyUI evicts models when VRAM gets tight, but free_memory() only considers models on the device it is freeing:

# comfy/model_management.py
if device is None or shift_model.device == device:

A text encoder loaded to the CPU is never on that device, so it is never a candidate. MiniMax H3's Qwen3-VL encoder is 15.7 GB. It runs once, produces the conditioning, and then sits in RAM for the entire render while the 21 GB diffusion model tries to find room. Peak demand was about 36 GB against 26 available.

--cache-none does not fix this. That flag clears the node output cache, which is a different structure from current_loaded_models.

The same pattern shows up again after sampling. The diffusion model is still staged when the VAE needs to decode, ComfyUI logs 0 models unloaded, and the VAE loads on top of it. On long clips that is where the OOM lands.

The nodes

Unload CPU Models frees everything loaded to a device other than the render device. In practice that is the text encoder. Wire it into the conditioning link between your encoder and your sampler.

Unload After Sampling frees everything except the VAE. Wire it into the latent link between your sampler and your VAE decode. This cannot affect your output: the latents are already computed, so the weights that produced them are dead memory by the time the node runs.

Both take any type in and return it unchanged, so you can drop them into an existing link without rewiring anything else.

Both have an optional keep field, a comma-separated list of substrings matched against model class names. Unload After Sampling defaults to VAE. Leave these alone unless something you still need is being unloaded.

Install

Via ComfyUI-Manager, search for "Unload Models".

Or clone it:

cd ComfyUI/custom_nodes
git clone https://github.com/neezoy/ComfyUI-UnloadModels

No dependencies beyond ComfyUI itself. Restart ComfyUI afterwards.

When this helps

Any workflow where a large model is used once early and then idles while something larger loads. Video models are the obvious case because the encoders are big, but nothing here is specific to video, to MiniMax H3, or to AMD.

If your models comfortably fit in VRAM, these nodes will not speed anything up. The gain comes from avoiding transfers across PCIe, so it scales with how badly you are over budget.

Notes

Set the log level to DEBUG to see each model considered and the reason it was kept or unloaded. At INFO you get one line per node with the total freed.

Other node packs include VRAM cleanup buttons. Those generally call ComfyUI's own free_memory(), which is exactly the function that skips off-device models. These nodes bypass it and unload directly.