ModelPack CLIP
The Text Encoder Loader That Won't Go Stale
- CLIP
By 2026 the text encoder stopped being an afterthought. Flux wants CLIP-L plus T5-XXL, SD3 wants three encoders, the video and audio models each want their own, and the new LLM-based encoders (Qwen3, Mistral) are big enough that they're often the memory bottleneck rather than the diffusion model. ModelPack CLIP is a text-encoder loader for that world - and it's the one node in this pack that does something core's loader can't.
The interesting part
Most custom CLIP loaders hardcode a list of architectures. That list rots: you write it for the ComfyUI of today, and every model generation after that needs a code change. This node builds its type dropdown at load time from the running ComfyUI's own comfy.sd.CLIPType enum:
inputs["required"]["type"] = ([t.name.lower() for t in comfy.sd.CLIPType],)
So the choices aren't the author's list, they're your build's list - around 35 entries on a current install, named stable_diffusion, stable_cascade, sd3, stable_audio, hunyuan_dit, flux, mochi, ltxv, hunyuan_video, pixart, cosmos, lumina2 and a couple dozen more. Update ComfyUI and the node's dropdown grows with it. You never get "this loader doesn't know what a Krea 2 encoder is yet."
Inputs and the output
reference- the OCI artifact reference, e.g.registry.example.com/models/my-clip:v1. Empty means local mode.file- with areference, the path inside the artifact; without one, a filename frommodels/text_encoders. Exactly one of the two paths gets taken.type- pick what the model card says. This is the field that decides how the file's weights get assembled into a conditioning encoder, and it is not guessable: a T5 file loaded asstable_diffusionis a silent mess.device(optional, advanced) -defaultorcpu.
CLIP is the only output. It goes into a CLIP Text Encode node - the architecture-flavoured one if your model needs it, CLIPTextEncodeFlux being the usual example.
That device field is the sleeper feature. Set it to cpu and the encoder loads on the CPU instead of the GPU. On an LLM-encoder model where the text encoder is eating 10+ GB of VRAM you'd rather spend on the diffusion model, parking it on CPU is a legitimate trade: slower prompt encoding, more room to sample. It's the same trick as the KB's note that the encoder is the component you're free to squeeze independently of the model.
Installing it
Manager: search ComfyUI ModelPack (cerussite). Or:
cd ComfyUI/custom_nodes
git clone https://github.com/SiLeader/ComfyUI-ModelPack comfyui-modelpack
python -m pip install -r comfyui-modelpack/requirements.txt
Restart afterwards. The dependency is just modelpack-client plus oras, jsonschema and zstandard - nothing that touches your torch install. Use the Python that runs ComfyUI (portable installs: python_embedded\python.exe), and note the pack needs Python 3.10+ and ComfyUI 0.22.0+.
The trap in the type field
Here's the bit the tooltip won't tell you. The node resolves your chosen type with:
clip_type = getattr(comfy.sd.CLIPType, type.upper(), comfy.sd.CLIPType.STABLE_DIFFUSION)
That trailing argument is a fallback. If the type name isn't in your build's enum, you don't get an error - you get stable_diffusion. In the UI you can't type a bad name, so this only bites when a workflow was saved on a newer ComfyUI than yours: the widget carries a value your build has never heard of, and the node quietly loads your SD3 encoder as if it were SD 1.5. The symptom is bad output, not a crash, which is exactly the kind of bug that eats an evening. If a shared workflow's text encoding looks wrong and the node looks fine, check whether its type value exists in your ComfyUI.
Pull behaviour, quickly
Weights land in ComfyUI/models/modelpack/text_encoders/<hash>/ with the hash keyed on the reference string. Digest references (@sha256:…) are pulled once and reused without touching the registry again; tag references are re-pulled on first use after each restart, falling back to your existing copy if the registry is unreachable. Private registries use docker login credentials from ~/.docker/config.json.
Two practical notes. The node hands ComfyUI exactly one file path, so a dual-encoder model whose artifact keeps CLIP-L and T5 as separate files needs you to check what's really inside before choosing type - leave file blank on a multi-weight artifact and the error message lists every weight path it found, which is the fastest inventory tool you have. And it only considers files with extensions ComfyUI's loader accepts; a .gguf weight only appears if you have the GGUF node pack installed, since that's what registers the extension.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| reference | STRING | OCI reference, e.g. registry.example.com/models/foo:v1. Leave empty to select a local file. | |
| file | STRING | Path within the artifact if it has multiple weights, or local ComfyUI model filename. | |
| type | COMBO | 35 options: stable_diffusion, stable_cascade, sd3, stable_audio, hunyuan_dit, flux, +29 | |
| deviceopt | COMBO | 2 options: default, cpu |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| CLIP | CLIP | — |