RDNA35 Generic PISA Benchmark
The four-way attention shootout — PISA vs SDPA vs Flash vs SageAttention on your ROCm card
- benchmark
Every AMD user's favorite question: is my attention path the bottleneck, and would anything else be faster? RDNA35 Generic PISA Benchmark answers it with a four-way kernel shootout on one synthetic input: generic PISA, ComfyUI's PyTorch SDPA path, Flash Attention, and - if you've installed it - SageAttention's ROCm7 build. You drop it in, pick a shape, and get back a text report of GPU-event medians plus a cosine-similarity check against dense SDPA for each approximate contender.
What it does
It fabricates Q/K/V at your chosen shape and runs each backend on identical tensors, reporting median times, the PISA/competitor ratio, and cosine similarity + mean abs error vs dense SDPA. The shape knobs are batch, heads, tokens (8192–32768, stepped by 64 - generic PISA here only handles T >= 8192), head_dim (32–256), exact_budget, and dtype (fp16 or bf16).
One mechanism detail worth knowing: the first-order correction path (the part that makes PISA's HYD routing score accurate) only activates for BF16 D=128 - that's where the fused CK statistics kernel lives. For FP16 and other head dims it falls back to a fused Triton statistics profile, and the benchmark floors your exact_budget at 0.25 to match. So if you benchmark FP16 D=64, you're measuring a different (cheaper, less accurate) PISA than the BF16 D=128 headline number.
Reading the output
The benchmark STRING lists all four backends. The useful bits:
- The PISA vs PyTorch SDPA ratio - for long sequences that's usually where PISA wins; at
T=8192, D=64the README shows it beating SDPA and roughly matching Flash. - The cosine vs dense SDPA line for PISA. This is the number that keeps you honest: on random inputs the pack's own numbers are ~0.70–0.81, i.e. generic PISA is visibly approximate. Kernel speed without quality validation is a footgun, not a feature.
- The SageAttention line - and the catch. SageAttention ROCm7 only supports its advertised head dims, and the node reports unsupported shapes as "unavailable" without aborting the rest of the benchmark. Installing SageAttention on ROCm is notoriously fiddly (a long-standing community pain point - the README's own changelog had to switch a workflow away from a CUDA-only SageAttention path), so expect that line to say "unavailable" more often than not.
Installing it
Manager → search RDNA35 Attention, or git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention into custom_nodes, restart. Base install is zero-dependency. SageAttention is optional and separate - the README's command is a --no-deps wheel install from the SageAttention-Rocm7 v1.0.6 release into ComfyUI's Python environment. You also need a ROCm PyTorch build on an RDNA3.5 card for the Triton-backed PISA path to run at all; without it the benchmark degrades to a two-way SDPA-vs-Flash comparison.
The honest verdict
This is a research workbench, not a "make my videos faster" node - nothing it measures changes your workflow. Its real value is deciding whether to bother with the pack's PISA patch on your model: if PISA loses to Flash at your actual shapes, you've saved yourself the quality risk of an approximate backend for nothing. If it wins and the cosine holds up at model level, you have a reason to try the Anima patch. Just never quote a PISA speedup without quoting its cosine, or you'll be the person who shipped a broken image to prove a point.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| batch | INT | 11–8 | — |
| heads | INT | 101–40 | — |
| tokens | INT | 81928192–32768 | — |
| head_dim | INT | 6432–256 | — |
| exact_budget | FLOAT | 0.25000.015625–1 | — |
| dtype | COMBO | float16 | 2 options: float16, bfloat16 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| benchmark | STRING | — |