ComfyUI Node
Attention Map Bending (Experimental)
Bends the cross-attention maps of a video DiT (WAN 2.1 / 2.2) frame by frame with any bending module (rotate, scale, translate, flip, blur, sharpen, multiply = amplify, ...), and self-attention through its outputs (self_query) or values (self_key). Experimental: validated only on tiny random-weight WAN models in the test suite, not yet on real video model weights.
Attention Map Bending (Experimental)
- model
- bending_module
- clip
- MODEL
- report
◄attentioncross_text►
◄blocks*►
◄tokensall►
◄renormalizekeys►
◄apply_toboth►
◄steps*►
◄t_start1.00►
◄t_end0.00►
◄heads*►
◄frames*►
◄strength1.00►
◄prompt►
◄strictfalse►
Categorymodel_bending/video (experimental)
Inputs (16)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| bending_module | BENDING_MODULE | — | |
| attention | COMBO | cross_text | cross_text: video-to-prompt cross-attention. cross_image: video-to-CLIP-image cross-attention (WAN 2.1 I2V). self_query: bend where each position's self-attention result lands. self_key: bend where each position reads from (values moved, layout kept). |
| blocks | STRING | * | Block indices, e.g. '13-18' (the middle of WAN 2.1 1.3B's 30 blocks; 14B has 40). '*' bends every block, the strongest effect. |
| tokens | STRING | all | cross attention only: 'all' (every key, padding included), 'prompt' (only the prompt's own tokens, not the padding), indices '0-3, 7', or prompt words 'horse, red car' (needs clip + prompt). Single words give much weaker effects than 'all'. |
| renormalize | COMBO | keys | keys: each video position's attention sums to 1 again (a valid softmax). per_token_mass: each token keeps its total attention. none: raw (amplify only acts without renormalisation). |
| apply_to | COMBO | both | CFG passes to bend. 'cond' alone is amplified by the guidance scale: stronger, glitchier. |
| stepsopt | STRING | * | Sampling steps, e.g. '0-2'. Early steps change layout and coherence; late steps change little. |
| t_startopt | FLOAT | 1.000–1 | — |
| t_endopt | FLOAT | 0.000–1 | — |
| headsopt | STRING | * | Attention heads, e.g. '0-3' (WAN 1.3B has 12, 14B has 40) |
| framesopt | STRING | * | Latent frames to bend, e.g. '0-5' (WAN: 1 latent frame = 4 video frames) |
| strengthopt | FLOAT | 1.00-4–4 | Blend between the unbent (0) and bent (1) map |
| clipopt | CLIP | Text encoder, to find prompt words given in 'tokens' | |
| promptopt | STRING | The positive prompt, to find words given in 'tokens' | |
| strictopt | BOOLEAN | false | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |
| report | STRING | — |