RDNA35 Patch PISA Attention
The Anima PISA patch — ~1.85x faster attention that will visibly change your images. Read this before you flip it on.
- model
- model
- info
Anima is Circlestone's 2B anime DiT on a Cosmos-Predict2 backbone, and it's slow - per-step compute is high for its size and it wants 30–50 steps. RDNA35 Patch PISA Attention is a model-local patch that keeps Anima's image self-attention in layers 0–19 dense and converts layers 20–27 to PISA, the piecewise-sparse attention scheme: only 23 of every 144 key/value blocks are computed exactly, the rest are routed by a cheap HYD score. On the author's gfx1151 stack that made the complete spatial attention call ~1.85x faster than ComfyUI's Flash path - a 45.9% reduction on the attention call itself.
But the README's layer-search table is the part you need to read before celebrating. PISA is approximate, and on Anima it changes composition. With the default 20–27 profile at matched seeds, SSIM against the dense Flash output landed around 0.72 / 0.63 across two seeds - visibly different images, judged better-looking by the author's aesthetic read, but not faithful reproductions. If same-seed fidelity matters more to you than aesthetics, the 24–27 profile is the conservative pick (SSIM ~0.91 / 0.78). The 4–11+20–27 schedule was outright rejected for text/layout regressions. This is a quality trade, not a free lunch.
What the node gives you
Inputs: model, enabled (on by default), verbose_fallbacks, plus the two research knobs - first_pisa_layer (default 20) and last_pisa_layer (default 27), both 0–27. One node represents one contiguous range; chaining nodes can create disjoint research schedules if you're feeling brave. Outputs: model (the cloned, patched model - wire it to your sampler) and info, a STRING that reports whether the direct PISA integration installed, on which layers, and at which exact-block count, or why it didn't (e.g. the wheel is missing or the model isn't the validated Anima profile).
The safety rails are strict, in keeping with the pack's design: it only patches Anima image self-attention modules at T=9216, H=16, D=128, BF16, contiguous, unmasked, forward-only. Cross-attention, masked calls, FP16 calls, and other sequence lengths chain to the previous ComfyUI backend. If a PISA call fails once it's started, failures are surfaced rather than silently retried on the same async stream. The spatial sparse budget is fixed at 23/144 - the only one validated; 32, 33, and 36 blocks all produced non-finite output in real generation.
Installing it
The pack itself is easy: Manager → search RDNA35 Attention, or git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention into custom_nodes, restart. The PISA patch is not. It needs the compiled rdna35_pisa_ck wheel from native/rdna35_pisa_ck, built against your exact PyTorch/ROCm runtime with MAX_JOBS=32 and installed with pip install --no-deps. The prebuilt wheel is Windows-only, BF16-only, gfx1151-only - so on Linux, or on FP16, or on any non-9070-series card, the node returns your model unchanged with the reason in info. Run the pack's Diagnostics node first and check torch.version.hip, your gfx target, and Triton before you even start building wheels.
The workflow you'll actually want
The README ships two benchmark workflows (pure BF16 and INT8 ConvRot) that store layers 20–27 and disable SaveImage so disk I/O stays out of the timing. That's the right shape: run a same-seed dense baseline, run the PISA profile, compare the two images and the timing. And when you do, pair this node with RDNA35 PISA Runtime Report at the end of the graph - it verifies the patch actually executed (8 self-attention calls per forward at the default profile, zero eligible fallbacks) instead of silently falling back and making your "benchmark" measure nothing. The progress bar is not a valid GPU timing source on async attention paths; that's not my opinion, it's in the README.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| enabled | BOOLEAN | true | — |
| verbose_fallbacks | BOOLEAN | false | — |
| first_pisa_layeropt | INT | 200–27 | — |
| last_pisa_layeropt | INT | 270–27 | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| info | STRING | — |