Nodes/ComfyUI Essentials/πŸ”§ Flux Attention Seeker
ComfyUI Node Runs on cloud

πŸ”§ Flux Attention Seeker

Per-layer attention tweaking for Flux's text encoders

By cubiqΒ·Created 3 years agoΒ·Updated about a year agoΒ· 1,155
πŸ”§ Flux Attention Seeker
  • clip
  • CLIP
β—„apply_to_querytrueβ–Ί
β—„apply_to_keytrueβ–Ί
β—„apply_to_valuetrueβ–Ί
β—„apply_to_outtrueβ–Ί
β—„clip_l_01.00β–Ί
β—„clip_l_11.00β–Ί
β—„clip_l_21.00β–Ί
β—„clip_l_31.00β–Ί
β—„clip_l_41.00β–Ί
β—„clip_l_51.00β–Ί
β—„clip_l_61.00β–Ί
β—„clip_l_71.00β–Ί
β—„clip_l_81.00β–Ί
β—„clip_l_91.00β–Ί
β—„clip_l_101.00β–Ί
β—„clip_l_111.00β–Ί
β—„t5xxl_01.00β–Ί
β—„t5xxl_11.00β–Ί
β—„t5xxl_21.00β–Ί
β—„t5xxl_31.00β–Ί
β—„t5xxl_41.00β–Ί
β—„t5xxl_51.00β–Ί
β—„t5xxl_61.00β–Ί
β—„t5xxl_71.00β–Ί
β—„t5xxl_81.00β–Ί
β—„t5xxl_91.00β–Ί
β—„t5xxl_101.00β–Ί
β—„t5xxl_111.00β–Ί
β—„t5xxl_121.00β–Ί
β—„t5xxl_131.00β–Ί
β—„t5xxl_141.00β–Ί
β—„t5xxl_151.00β–Ί
β—„t5xxl_161.00β–Ί
β—„t5xxl_171.00β–Ί
β—„t5xxl_181.00β–Ί
β—„t5xxl_191.00β–Ί
β—„t5xxl_201.00β–Ί
β—„t5xxl_211.00β–Ί
β—„t5xxl_221.00β–Ί
β—„t5xxl_231.00β–Ί

This is one of the more experimental nodes in the pack, and it helps to be honest about that up front: FluxAttentionSeeker lets you scale attention layer by layer inside Flux's two text encoders, and it's a research toy more than a daily driver. Flux reads your prompt through CLIP-L (12 layers) and T5-XXL (24 layers), and different layers capture different levels of meaning. This node gives you a slider per layer to turn each one's contribution up or down, then hands back a modified CLIP. Most people never touch it. The ones who do are hunting for a specific effect - better prompt adherence, a style shift - by "seeking" through the layers, which is where the name comes from.

It's from ComfyUI Essentials by cubiq (Matteo Spinelli, the ComfyUI_IPAdapter_plus author). Fitting that it comes from the IPAdapter dev - this is the same appetite for poking at attention internals, applied to Flux's text side.

How it works

Flux uses CLIP-L plus T5-XXL, and as the KB explains, "Flux takes your prompt or caption, and hands it to both T5 and CLIP." Each encoder is a stack of transformer layers, and attention is how each layer decides what in your prompt matters. This node multiplies the attention at each layer by a factor you set - leave a layer at 1.0 and it's untouched; push it up or down and you change how strongly that layer shapes the conditioning. You can also choose which parts of the attention mechanism get scaled (query, key, value, output). The result is a patched CLIP you feed to your encode step as usual.

The inputs and outputs that matter

  • clip (CLIP) - your Flux CLIP (the dual CLIP-L + T5 loader output).
  • clip_l_0 … clip_l_11 and t5xxl_0 … t5xxl_23 (FLOAT, 0–5, default 1) - per-layer scaling. 1 is neutral. This is the whole node; expect to sweep values, not to know them in advance.
  • apply_to_query / apply_to_key / apply_to_value / apply_to_out (BOOLEAN, default true) - which attention components to affect. Defaults are fine to start.
  • CLIP (output) - the patched encoder. Wire it into your text-encode node in place of the original CLIP.

How to install it

ComfyUI Manager: search ComfyUI Essentials, install, restart. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/cubiq/ComfyUI_essentials
pip install -r ComfyUI_essentials/requirements.txt

then restart. No extra models - it patches the Flux CLIP you already loaded.

Common issues & troubleshooting

Nothing seems to change. With every layer at 1.0 this node is a no-op - that's neutral. Effects only appear once you push layers away from 1, and even then the changes can be subtle. This is a "run a sweep and compare" tool, not a one-click switch.

It's a Flux-only node. The layer counts (12 for CLIP-L, 24 for T5) are Flux's dual-encoder setup. Don't expect it to do anything sensible on an SDXL or SD1.5 CLIP.

My results got worse / weird. Entirely possible - pushing attention scaling far from neutral can break coherence as easily as improve it. There's no universally good setting; this is exploratory. If you're not chasing a specific effect, leave it out of the graph.

Missing after a Comfy update. Essentials is maintenance-only since April 2025 and pack nodes have broken against newer ComfyUI. Update via Manager; roll back Comfy if it's genuinely incompatible.

Categoryessentials/conditioning

Inputs (41)

NameTypeDefaultDescription
clipCLIPβ€”
apply_to_queryBOOLEANtrueβ€”
apply_to_keyBOOLEANtrueβ€”
apply_to_valueBOOLEANtrueβ€”
apply_to_outBOOLEANtrueβ€”
clip_l_0FLOAT1.000–5β€”
clip_l_1FLOAT1.000–5β€”
clip_l_2FLOAT1.000–5β€”
clip_l_3FLOAT1.000–5β€”
clip_l_4FLOAT1.000–5β€”
clip_l_5FLOAT1.000–5β€”
clip_l_6FLOAT1.000–5β€”
clip_l_7FLOAT1.000–5β€”
clip_l_8FLOAT1.000–5β€”
clip_l_9FLOAT1.000–5β€”
clip_l_10FLOAT1.000–5β€”
clip_l_11FLOAT1.000–5β€”
t5xxl_0FLOAT1.000–5β€”
t5xxl_1FLOAT1.000–5β€”
t5xxl_2FLOAT1.000–5β€”
t5xxl_3FLOAT1.000–5β€”
t5xxl_4FLOAT1.000–5β€”
t5xxl_5FLOAT1.000–5β€”
t5xxl_6FLOAT1.000–5β€”
t5xxl_7FLOAT1.000–5β€”
t5xxl_8FLOAT1.000–5β€”
t5xxl_9FLOAT1.000–5β€”
t5xxl_10FLOAT1.000–5β€”
t5xxl_11FLOAT1.000–5β€”
t5xxl_12FLOAT1.000–5β€”
t5xxl_13FLOAT1.000–5β€”
t5xxl_14FLOAT1.000–5β€”
t5xxl_15FLOAT1.000–5β€”
t5xxl_16FLOAT1.000–5β€”
t5xxl_17FLOAT1.000–5β€”
t5xxl_18FLOAT1.000–5β€”
t5xxl_19FLOAT1.000–5β€”
t5xxl_20FLOAT1.000–5β€”
t5xxl_21FLOAT1.000–5β€”
t5xxl_22FLOAT1.000–5β€”
t5xxl_23FLOAT1.000–5β€”

Outputs (1)

NameTypeDescription
CLIPCLIPβ€”