ComfyUI Node
Attention Map Capture (Experimental)
Records the cross-attention maps of chosen prompt tokens (head mean, first conditional sample) for every step and chosen block while a sampler runs, after any upstream attention bends. Read them as video frames with 'Read Attention Maps'. Experimental: validated only on tiny random-weight WAN models in the test suite, not yet on real video model weights.
Attention Map Capture (Experimental)
- model
- clip
- MODEL
- attention_maps
- report
◄attentioncross_text►
◄blocks13-18►
◄tokensprompt►
◄steps*►
◄heads*►
◄prompt►
◄strictfalse►
Categorymodel_bending/video (experimental)
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| attention | COMBO | cross_text | 2 options: cross_text, cross_image |
| blocks | STRING | 13-18 | Blocks to record, e.g. '13-18' |
| tokens | STRING | prompt | Tokens to record: prompt words 'horse' (needs clip + prompt), indices '0-3', or 'prompt' (each prompt token, at most 32) |
| stepsopt | STRING | * | Sampling steps to record |
| headsopt | STRING | * | — |
| clipopt | CLIP | — | |
| promptopt | STRING | — | |
| strictopt | BOOLEAN | false | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |
| attention_maps | ATTENTION_MAPS | — |
| report | STRING | — |