ComfyUI Node
Read Attention Maps (Experimental)
Turns recorded attention maps into video frames (IMAGE batch): bright = the token attends there. Experimental: validated only on tiny random-weight WAN models in the test suite, not yet on real video model weights.
Read Attention Maps (Experimental)
- attention_maps
- latent
- frames
- report
◄tokensum►
◄blockmean►
◄stepall►
◄normalizeper_video►
◄colormapinferno►
◄match_video_framestrue►
Categorymodel_bending/video (experimental)
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| attention_maps | ATTENTION_MAPS | — | |
| latent | LATENT | Connect the output of the sampler that used the capture model, so this node runs after sampling | |
| token | STRING | sum | 'sum' of the recorded tokens, or a token index |
| block | STRING | mean | 'mean' over recorded blocks, or a block index |
| step | STRING | all | 'all' = every recorded step side by side, 'mean', or a step index |
| normalize | COMBO | per_video | 2 options: per_video, per_frame |
| colormap | COMBO | inferno | 2 options: inferno, gray |
| match_video_frames | BOOLEAN | true | Repeat each latent frame 4x (after the first) so the maps line up with WAN's decoded frames |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| frames | IMAGE | — |
| report | STRING | — |