Nodes/ComfyUI-Model-Bending/Attention Map Bending (Experimental)
ComfyUI Node

Attention Map Bending (Experimental)

Bends the cross-attention maps of a video DiT (WAN 2.1 / 2.2) frame by frame with any bending module (rotate, scale, translate, flip, blur, sharpen, multiply = amplify, ...), and self-attention through its outputs (self_query) or values (self_key). Experimental: validated only on tiny random-weight WAN models in the test suite, not yet on real video model weights.

By abuzreq·Created 2 years ago·Updated 2 days ago· 23
Attention Map Bending (Experimental)
  • model
  • bending_module
  • clip
  • MODEL
  • report
◄attentioncross_text►
◄blocks*►
◄tokensall►
◄renormalizekeys►
◄apply_toboth►
◄steps*►
◄t_start1.00►
◄t_end0.00►
◄heads*►
◄frames*►
◄strength1.00►
◄prompt►
◄strictfalse►
Categorymodel_bending/video (experimental)

Inputs (16)

NameTypeDefaultDescription
modelMODEL—
bending_moduleBENDING_MODULE—
attentionCOMBOcross_textcross_text: video-to-prompt cross-attention. cross_image: video-to-CLIP-image cross-attention (WAN 2.1 I2V). self_query: bend where each position's self-attention result lands. self_key: bend where each position reads from (values moved, layout kept).
blocksSTRING*Block indices, e.g. '13-18' (the middle of WAN 2.1 1.3B's 30 blocks; 14B has 40). '*' bends every block, the strongest effect.
tokensSTRINGallcross attention only: 'all' (every key, padding included), 'prompt' (only the prompt's own tokens, not the padding), indices '0-3, 7', or prompt words 'horse, red car' (needs clip + prompt). Single words give much weaker effects than 'all'.
renormalizeCOMBOkeyskeys: each video position's attention sums to 1 again (a valid softmax). per_token_mass: each token keeps its total attention. none: raw (amplify only acts without renormalisation).
apply_toCOMBObothCFG passes to bend. 'cond' alone is amplified by the guidance scale: stronger, glitchier.
stepsoptSTRING*Sampling steps, e.g. '0-2'. Early steps change layout and coherence; late steps change little.
t_startoptFLOAT1.000–1—
t_endoptFLOAT0.000–1—
headsoptSTRING*Attention heads, e.g. '0-3' (WAN 1.3B has 12, 14B has 40)
framesoptSTRING*Latent frames to bend, e.g. '0-5' (WAN: 1 latent frame = 4 video frames)
strengthoptFLOAT1.00-4–4Blend between the unbent (0) and bent (1) map
clipoptCLIPText encoder, to find prompt words given in 'tokens'
promptoptSTRINGThe positive prompt, to find words given in 'tokens'
strictoptBOOLEANfalse—

Outputs (2)

NameTypeDescription
MODELMODEL—
reportSTRING—