RDNA35 Exact Full Attention Benchmark
The exact-attention Triton kernel — a research benchmark for gfx1151, not a shortcut
- benchmark
Most attention research in the ComfyUI world is about skipping work - sparse attention, flash kernels, token pruning. RDNA35 Exact Full Attention Benchmark goes the other way: it's a hand-tuned Triton kernel that computes exact full attention with online softmax, and this node's only job is to tell you whether it's faster than PyTorch SDPA on your AMD card. That's it. It doesn't patch anything, doesn't touch your model, doesn't promise you a speedup - it measures one and reports the other.
Why it exists
The README is blunt: this is an "isolated gfx1151 research path," and none of it replaces normal ComfyUI attention globally. The author is exploring whether a purpose-built online-softmax Triton kernel can beat the generic SDPA path on RDNA3.5 hardware, where flash-attention support has historically been thinner than on NVIDIA. The kernel operates on [BH,Q,D] x [BH,K,D] tensors - the flattened batch-head layout - which is why the benchmark's shapes look like Anima's (the pack's other obsession is Anima's self- and cross-attention, and D=128 is Anima's head dim).
Mechanically it's an online-softmax Triton kernel: it iterates key blocks, keeps running maxes and exp-sums, normalizes at the end - the standard trick to avoid materializing the full QK^T matrix. The config is tuned per head dim (block_m/block_n/num_warps) and even per target, with waves_per_eu set to 1 on gfx1151 specifically. On any other GPU it just uses the generic config.
The inputs
All shape knobs, none of them scary:
query_tokensandkey_tokens- the Q and K sequence lengths (up to 9216, stepped by 64), so you can benchmark self-attention (equal) or cross-attention (different) shapes.head_dim- 64 or 128, the only two the kernel supports.dtype- float16 or bfloat16.batchandheads- the size of the flattenedBHdimension.
It creates random Q/K/V on the device, computes a dense SDPA reference, runs the Triton kernel, times both with GPU events, and returns one benchmark STRING: SDPA median, Triton median, the Triton/SDPA ratio, max abs error against the reference, and the kernel config that ran. A ratio under 1.0x means the custom kernel lost - the node will happily tell you that.
Installing and running it
Same pack, same story: ComfyUI Manager → search RDNA35 Attention, or
cd ComfyUI/custom_nodes
git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention
restart, and you'll find it under RDNA35/Attention Research. It's flagged experimental in the source. It hard-requires a ROCm GPU - the code allocates on cuda (which is what PyTorch ROCm calls its device type) and has no CPU fallback, so on an NVIDIA card or a CPU-only box it will simply fail. It also needs Triton to even consider running the kernel; without it you get a validation failure and no timing.
What to actually do with it
If you're not doing attention kernel research on RDNA3.5, honestly? Probably nothing - this is a measurement tool for the author's experiments, and its README position is "isolated research path." The one genuinely useful habit it supports: if you're building an Anima workflow on an RX 9070-series card and wondering whether ComfyUI's default SDPA is the thing slowing you down, this node gives you a clean shape-matched comparison. If the ratio comes back at or above 1x, move on - the bottleneck is elsewhere, and the pack's other nodes (the MiniMax-H3 layout patch) are where the real, measured wins live anyway.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| batch | INT | 11–4 | — |
| heads | INT | 161–40 | — |
| query_tokens | INT | 409664–9216 | — |
| key_tokens | INT | 409664–9216 | — |
| head_dim | COMBO | 128 | 2 options: 64, 128 |
| dtype | COMBO | bfloat16 | 2 options: float16, bfloat16 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| benchmark | STRING | — |