CCIP Adapter Infer
Where reference-image tokens actually get made
- adapter
- features
- out_hidden
- out_mask
CCIP Adapter Infer is the half of this pack that actually does math. Its loader sibling just reads a checkpoint off disk; this node takes those weights, runs them over a reference image's feature vector, and produces the token embeddings that make Lumina Image 2.0 "see" your character. It's the bridge between an anime character-identity feature and Gemma 2B's text-encoder space - which is why everything else in the workflow quietly depends on it.
The wiring in a real workflow is short: Load Image → the feature extractor from the sibling comfyui-ccip pack (CCIPExtractFeature) → this node → the packer/concatenation nodes that fold the tokens into your positive conditioning → KSampler. The example workflow shipped with the pack (ccip_adapter_lu2.json) does exactly that on a Lumina Image 2.0 graph, reference image in, same-character image out.
How it works, straight from the source: the node takes your feature tensor, pushes it to the GPU as float32, and runs the adapter under torch.no_grad(). One reference image becomes (1, K, out_dim) - K tokens of 2304 wide, which is Gemma 2B's hidden size - then both outputs are detached and moved back to CPU before returning. All of that is invisible to you; you just get two tensors.
The inputs that matter:
- adapter - the
CCIP_ADAPTERhandle from CCIP Adapter Loader. No handle, no run; it's the first thing that errors. - features - a
TENSORfrom the feature extractor, shape(N, in_dim). Its last dimension must equal the loader'sin_dimor the node raises.
Outputs:
- out_hidden -
(B, K, out_dim)float32. This is the payload. Wire it intoConditioningPacker(fromcomfyui-spawner-nodes) or a conditioning-concatenation node so the tokens land in the positive prompt. - out_mask -
(B, K)int32, and here's the honest bit: it's all ones, every time. It's a placeholder so downstream nodes that expect a token mask have something to read. Don't read anything into it.
The trap with this node is that it's silent. It's not an output node, it just returns tensors, so if you forget to wire out_hidden into the conditioning you get no error and no character consistency - just a prompt that ignores your reference image entirely. The workflow runs fine and the difference is in the output. Second-most common burn: the feature dimension mismatch, usually because someone changed in_dim on the loader after the extractor was already configured. Third, worth knowing: because the loader loads its checkpoint with strict=True, a mismatch between the file's shape and the loader's numbers surfaces upstream in the loader, not here - so a failed run often points you at the wrong node.
Install is the same chain as the loader, because they're one pack. Make sure the two prerequisites are present first:
cd ComfyUI/custom_nodes
git clone https://github.com/spawner1145/comfyui-ccip-adapter
then restart. The feature extractor's comfyui-ccip dependency needs onnxruntime or onnxruntime-gpu and may download its own model weights on first use; the adapter checkpoint goes in ComfyUI/models/ccip_adapter/. Heavy, but the whole thing is small - an adapter that fits in a single file, doing one job, and when you've got your character consistency it's easy to forget it's even there.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| adapter | CCIP_ADAPTER | — | |
| features | TENSOR | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| out_hidden | TENSOR | — |
| out_mask | TENSOR | — |