CCIP Adapter Loader
The boring node that makes 'keep this character' work in Lumina Image 2.0
- adapter
CCIP Adapter Loader doesn't touch a single pixel, and that's the point. It's the loading half of a two-node reference-image adapter for Lumina Image 2.0: it reads a small adapter checkpoint from disk and hands the weights to its sibling, CCIP Adapter Infer, wrapped in a CCIP_ADAPTER handle. If you pulled a workflow that contains one, your whole job is usually two steps - drop the .safetensors in the right folder, leave the three numbers alone - and it quietly works.
What's actually going on here is more interesting than the node looks. Lumina Image 2.0 conditions on text through Gemma 2B (the workflow's CLIPLoader points at gemma_2_2b_fp16.safetensors with type lumina2). This adapter is the bridge that lets you feed it a reference image instead of only a sentence: a character-identity feature vector gets projected into Gemma's token space as a stack of prompt tokens. The default out_dim of 2304 isn't random - that's Gemma 2B's hidden width. The default tokens_per_ref of 32 is how many token vectors one reference image becomes. Push those into the conditioning and the model "sees" your character in the same breath as the text, which is the trick behind same-girl-different-pose workflows on this stack.
Mechanically, per the source, the loader builds a small MLP (CCIPToGemmaAdapter): a base projection (in_dim → out_dim, GELU, out_dim → out_dim), a learned per-token embedding (tokens_per_ref × out_dim), then a LayerNorm plus a two-layer MLP whose final layer is zero-initialized so the adapter starts life as a near-identity residual. It loads your checkpoint with strict=True and flips to eval mode. Translation: those three integers define the network architecture, not just metadata. The checkpoint must match them exactly or the load fails.
The inputs that actually matter:
- adapter_name - dropdown populated from
ComfyUI/models/ccip_adapter/. If that folder is empty you'll see(no adapter found in models/ccip_adapter)and the node throws a clear error telling you where to put the file. - in_dim (default 768) - the feature width coming in from the extractor node upstream. Match it to what your extractor emits.
- out_dim (default 2304) - Gemma 2B's hidden size. Leave it unless your checkpoint was trained differently.
- tokens_per_ref (default 32) - tokens minted per reference image.
It returns a single adapter output, wired straight into CCIP Adapter Infer. That's the whole node.
Install is the fiddly part, because this pack is the last in a chain. The README is explicit: install comfyui-ccip and comfyui-spawner-nodes first, then:
cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-ccip-adapter
or just search "ccip-adapter" in ComfyUI Manager and restart. This pack has no pip requirements of its own - torch and safetensors are already in ComfyUI. The sibling comfyui-ccip pack needs onnxruntime (or onnxruntime-gpu) for its feature extractor, and it may pull model weights from Hugging Face on first run.
Where people get burned: forgetting the checkpoint entirely (empty-folder error), or editing the three INTs to "match" a file that was trained with defaults. The strict load means guessing wrong is a hard failure, not a subtly bad image. And after you drop a new .safetensors into models/ccip_adapter/, the dropdown only refreshes on restart. Load the file, wire the handle into the infer node, and the fun half of the pair takes over.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| in_dim | INT | 7681–16384 | — |
| out_dim | INT | 23041–16384 | — |
| tokens_per_ref | INT | 321–4096 | — |
| adapter_name | COMBO | (no adapter found in models/ccip_adapter) | 1 options: (no adapter found in models/ccip_adapter) |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| adapter | CCIP_ADAPTER | — |