GGUF UNET Loader
The 12GB lever, plus two knobs you can ignore
- MODEL
GGUF is how most people actually run Flux-class and video models on consumer cards, and this is the pack's version of the loading node. If you already run city96's ComfyUI-GGUF, you have this covered. If you'd rather install one utility pack than three, this one is bundled - the pack's own source headers credit city96's Apache-2.0 code, and py/gguf_core/ here is that code (loader, ops, dequant routines) with a wrapper around it.
What it's for
The Q ladder hasn't really changed: Q8 is essentially fp16 at half the size, Q6 shows minimal visible difference, Q5_K is the last tier with negligible loss, Q4_K_M is the accepted 12GB-card compromise, and Q3/Q2 are for the genuinely stuck. Counterintuitively, Q8 is often faster than Q2/Q3 - 4- and 8-bit quants get hardware dot-product instructions, lower ones pay a dequantization penalty.
The trap worth repeating: reach for quantization when the model doesn't fit, not by default. A 12GB card holding an fp8 Flux 2 Klein gains almost nothing from Q5 GGUF, and at least one A/B found the real 2x speedup came from removing --lowvram --reserve-vram flags rather than quantizing at all.
Inputs and outputs
unet_name is the only required input, and it lists .gguf files only. Optional:
dequant_dtype-default,target,float32,float16,bfloat16patch_dtype- same five options
Output is a single MODEL, ready for your sampler, guider, or a model-patching chain.
What do the dtype knobs do? default leaves the quantization ops at their own choice. target means "dequantize to the dtype the model targets". The rest force a specific torch dtype for dequantized weights and for LoRA patch computation respectively. Honest advice: leave both on default. float32 is a VRAM bonfire, and bfloat16 is the only one worth an experiment on a modern card. These are the same knobs as city96's advanced loader, and most people never touch them.
How it works
The node registers a unet_gguf file type pointing at your unet / diffusion_models folders with a .gguf extension filter, reads the file through gguf_sd_loader, then calls ComfyUI's load_diffusion_model_state_dict with the pack's GGMLOps as custom operations. The result is wrapped in the pack's GGUFModelPatcher, whose whole job is surviving LoRA patches on quantized weights.
That wrapper is also the reason the standard warning applies here: applying a LoRA to a GGUF model is slower, because each layer has to be dequantized, patched, and re-quantized as it runs. If you're VRAM-capped and want LoRAs, dropping a quant level to make room usually beats staying at Q8.
Install and model placement
cd ComfyUI/custom_nodes
git clone https://github.com/wenchengxiang/ComfyUI-Practical-Tools.git
gguf is in the pack's requirements.txt; without it the GGUF modules can't import. Files go here:
ComfyUI/models/diffusion_models/ # or models/unet/ - same folder list
ComfyUI/models/text_encoders/ # for the CLIP side, if you're doing that too
Common issues
The dropdown is empty. Almost always folder placement - this is the single most common ComfyUI complaint there is, and it's usually not a bug. Only .gguf files show in this node; a safetensors diffusion model won't appear here and needs the regular "Load Diffusion Model" path. After copying files, refresh or restart: the file list is cached and only re-checked when a folder's mtime changes.
"ERROR: Could not detect model type of: …" - the file loaded but ComfyUI couldn't identify the architecture. Either it isn't a diffusion transformer at all, or it's a quant or architecture the bundled code doesn't know. The dequant module handles BF16, Q8_0, Q6_K, Q5_K/Q5_1/Q5_0, Q4_K/Q4_1/Q4_0, Q3_K, Q2_K, IQ4_NL and IQ4_XS; anything outside that list falls back to a numpy path the source itself labels "incredibly slow".
The node isn't in the menu at all. The pack's __init__.py loads every node file in its own try/except and prints a [WCX Nodes Error] line to the console when one fails. If gguf didn't install (or a pip step was skipped on an offline box), you'll get a fully working pack minus the loaders. Check the startup console, not the node search.
You have two GGUF loaders now. If city96's pack is also installed, both work; they're separate node classes. Pick one and stay with it inside a workflow so you're not debugging two dequantization paths at once.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| unet_name | COMBO | 0 options: | |
| dequant_dtypeopt | COMBO | default | 5 options: default, target, float32, float16, bfloat16 |
| patch_dtypeopt | COMBO | default | 5 options: default, target, float32, float16, bfloat16 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |