GGUF CLIP Loader
The half of the stack you forgot to quantize
- CLIP
On the models that took over in late 2025 - Flux 2, Qwen-Image, Z-Image - the text encoder is a separate file with its own VRAM budget, and it's frequently what decides whether a model fits at all. A 24B Mistral encoder at fp16 is roughly 18GB before you've loaded any diffusion weights. Quantizing the encoder hard and keeping precision in the diffusion model is one of the better-understood levers in current ComfyUI, and this node is the pack's version of that.
What it is
It's ComfyUI's core CLIPLoader shape - clip_name plus type - with a GGUF code path bolted in. The GGUF machinery is the pack's bundled copy of city96's ComfyUI-GGUF (py/gguf_core/, Apache-2.0, credited in the file headers), so the mechanism is the familiar one: read the quantized file, register the GGML quantization ops as custom operations for ComfyUI's model loading, and build the text encoder from the resulting state dict.
Concretely: load_text_encoder_state_dicts is called with GGMLOps as custom operations and the text-encoder offload device as the initial device - so there's no device widget here like the core node has, it just starts offloaded. The patcher is then re-wrapped in the pack's GGUFModelPatcher, which is what keeps LoRA patching from blowing up on quantized weights.
Inputs and outputs
clip_name(required) - a list that merges regular CLIP files and.gguffiles, so you can point this at a normal safetensors encoder too. It's effectively a drop-in for the core loader.type(required) - taken straight from your ComfyUI's ownCLIPLoader, so it tracks whatever core supports:stable_diffusion,sd3,wan,ltxv,hidream,chroma,lumina2,qwen_image,flux2and about twenty more.
Output is one CLIP, which goes into your CLIP Text Encode node (or the model's instruction-encoder node) exactly as the core loader's output would.
The type is the field beginners get wrong and it's worth reading the core node's own notes for the recipe: sd3 wants t5xxl / clip-g / clip-l, wan wants umt5-xxl, hidream takes llama-3.1 (recommended) or t5, stable_diffusion is clip-l. Pick the wrong one and you get weird conditioning or a load failure with no obvious cause.
Install and where the files go
Manager → search Practical-Tools, or:
cd ComfyUI/custom_nodes
git clone https://github.com/wenchengxiang/ComfyUI-Practical-Tools.git
gguf is in the pack's requirements.txt and is required for this node to import. Text encoders - GGUF or not - live in:
ComfyUI/models/text_encoders/ # legacy alias: models/clip/
Common issues
It refuses scaled FP8. If a file contains scaled-fp8 tensors, the loader raises rather than guess: "Mixing scaled FP8 with GGUF is not supported! Use regular CLIP loader or switch model(s)". That check fires on any non-GGUF file passed in here, so if you're loading a scaled-fp8 encoder - including by itself - use the core CLIP loader instead. This node is for .gguf.
The node is missing from the menu entirely. The pack's __init__.py loads each node file inside its own try/except and prints a [WCX Nodes Error] line when one fails. If the gguf Python package isn't installed, the loader file dies on import and you get the rest of the pack working perfectly, with no GGUF loaders anywhere. Look at the startup console.
Empty dropdown, again. Nothing in models/text_encoders. Restart or refresh after copying files - the list is cached and only re-checked when a folder's mtime changes. And keep the encoder's quant level in perspective: a Q5 or Q8 text encoder is widely reported as fine, though the quality cost of going lower to Q4 is contested, and on weaker cards it's very often the difference between a model loading and not.
One naming caveat: if a type string isn't recognised it silently falls back to stable_diffusion rather than erroring. If a model that should work is producing mushy output, check what you picked for type before you start blaming the quant level.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| clip_name | COMBO | 0 options: | |
| type | COMBO | 28 options: stable_diffusion, stable_cascade, sd3, stable_audio, mochi, ltxv, +22 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| CLIP | CLIP | — |