Nodes/ComfyUI ModelPack/ModelPack CLIP
ComfyUI Node

ModelPack CLIP

The Text Encoder Loader That Won't Go Stale

By SiLeader·Created 24 days ago·Updated 21 days ago· 1
ModelPack CLIP
    • CLIP
    ◄reference►
    ◄file►
    ◄type▾►
    ◄device▾►

    By 2026 the text encoder stopped being an afterthought. Flux wants CLIP-L plus T5-XXL, SD3 wants three encoders, the video and audio models each want their own, and the new LLM-based encoders (Qwen3, Mistral) are big enough that they're often the memory bottleneck rather than the diffusion model. ModelPack CLIP is a text-encoder loader for that world - and it's the one node in this pack that does something core's loader can't.

    The interesting part

    Most custom CLIP loaders hardcode a list of architectures. That list rots: you write it for the ComfyUI of today, and every model generation after that needs a code change. This node builds its type dropdown at load time from the running ComfyUI's own comfy.sd.CLIPType enum:

    inputs["required"]["type"] = ([t.name.lower() for t in comfy.sd.CLIPType],)
    

    So the choices aren't the author's list, they're your build's list - around 35 entries on a current install, named stable_diffusion, stable_cascade, sd3, stable_audio, hunyuan_dit, flux, mochi, ltxv, hunyuan_video, pixart, cosmos, lumina2 and a couple dozen more. Update ComfyUI and the node's dropdown grows with it. You never get "this loader doesn't know what a Krea 2 encoder is yet."

    Inputs and the output

    • reference - the OCI artifact reference, e.g. registry.example.com/models/my-clip:v1. Empty means local mode.
    • file - with a reference, the path inside the artifact; without one, a filename from models/text_encoders. Exactly one of the two paths gets taken.
    • type - pick what the model card says. This is the field that decides how the file's weights get assembled into a conditioning encoder, and it is not guessable: a T5 file loaded as stable_diffusion is a silent mess.
    • device (optional, advanced) - default or cpu.

    CLIP is the only output. It goes into a CLIP Text Encode node - the architecture-flavoured one if your model needs it, CLIPTextEncodeFlux being the usual example.

    That device field is the sleeper feature. Set it to cpu and the encoder loads on the CPU instead of the GPU. On an LLM-encoder model where the text encoder is eating 10+ GB of VRAM you'd rather spend on the diffusion model, parking it on CPU is a legitimate trade: slower prompt encoding, more room to sample. It's the same trick as the KB's note that the encoder is the component you're free to squeeze independently of the model.

    Installing it

    Manager: search ComfyUI ModelPack (cerussite). Or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/SiLeader/ComfyUI-ModelPack comfyui-modelpack
    python -m pip install -r comfyui-modelpack/requirements.txt
    

    Restart afterwards. The dependency is just modelpack-client plus oras, jsonschema and zstandard - nothing that touches your torch install. Use the Python that runs ComfyUI (portable installs: python_embedded\python.exe), and note the pack needs Python 3.10+ and ComfyUI 0.22.0+.

    The trap in the type field

    Here's the bit the tooltip won't tell you. The node resolves your chosen type with:

    clip_type = getattr(comfy.sd.CLIPType, type.upper(), comfy.sd.CLIPType.STABLE_DIFFUSION)
    

    That trailing argument is a fallback. If the type name isn't in your build's enum, you don't get an error - you get stable_diffusion. In the UI you can't type a bad name, so this only bites when a workflow was saved on a newer ComfyUI than yours: the widget carries a value your build has never heard of, and the node quietly loads your SD3 encoder as if it were SD 1.5. The symptom is bad output, not a crash, which is exactly the kind of bug that eats an evening. If a shared workflow's text encoding looks wrong and the node looks fine, check whether its type value exists in your ComfyUI.

    Pull behaviour, quickly

    Weights land in ComfyUI/models/modelpack/text_encoders/<hash>/ with the hash keyed on the reference string. Digest references (@sha256:…) are pulled once and reused without touching the registry again; tag references are re-pulled on first use after each restart, falling back to your existing copy if the registry is unreachable. Private registries use docker login credentials from ~/.docker/config.json.

    Two practical notes. The node hands ComfyUI exactly one file path, so a dual-encoder model whose artifact keeps CLIP-L and T5 as separate files needs you to check what's really inside before choosing type - leave file blank on a multi-weight artifact and the error message lists every weight path it found, which is the fastest inventory tool you have. And it only considers files with extensions ComfyUI's loader accepts; a .gguf weight only appears if you have the GGUF node pack installed, since that's what registers the extension.

    CategoryModelPack/loaders

    Inputs (4)

    NameTypeDefaultDescription
    referenceSTRINGOCI reference, e.g. registry.example.com/models/foo:v1. Leave empty to select a local file.
    fileSTRINGPath within the artifact if it has multiple weights, or local ComfyUI model filename.
    typeCOMBO35 options: stable_diffusion, stable_cascade, sd3, stable_audio, hunyuan_dit, flux, +29
    deviceoptCOMBO2 options: default, cpu

    Outputs (1)

    NameTypeDescription
    CLIPCLIP—