Nodes/ComfyUI RDNA35 Attention/RDNA35 PISA Attention Benchmark
ComfyUI Node

RDNA35 PISA Attention Benchmark

PISA vs dense SDPA, measured honestly — including the compile tax everyone forgets

By Yasei-no-otoko·Created 3 months ago·Updated 2 months ago· 1
RDNA35 PISA Attention Benchmark
    • benchmark
    ◄batch_heads1►
    ◄tokens9216►
    ◄exact_budget0.1563►
    ◄dtypebfloat16►
    ◄backendck►

    PISA - Piecewise Sparse Attention - is the hot 2026 idea that most attention is wasted on uninformative key/value blocks, so you keep a fraction of blocks exact and approximate the rest. It's real, it's in production models, and on paper it's a huge win. The catch is that every PISA benchmark you'll see online conveniently forgets the first-run cost. RDNA35 PISA Attention Benchmark is the node that doesn't: it reports the cold compile-and-call time and the steady-state time as separate numbers, so you can see exactly what you're signing up for.

    What it measures

    It builds synthetic [BH, T, 128] tensors and runs the pack's PISA path - a hybrid of a compiled CK Tile statistics kernel (for the HYD routing score), PyTorch FlexAttention for the exact blocks, and a WMMA first-order correction - against dense SDPA. The report splits:

    • cold compile/call time (first use, when FlexAttention and the centroid approximation compile through Inductor/Triton),
    • steady-state GPU-event median,
    • dense SDPA median, the PISA/dense ratio, cosine similarity and mean abs error vs dense, and the fallback reason if the hybrid didn't run.

    It also tells you how many blocks stayed exact, e.g. exact blocks: 23/144 at the default exact_budget of 0.15625. That default matters: 23/144 is the only spatial profile validated on real Anima models in this pack - budgets at 32/33/36 went non-finite during actual generation. More sparsity is not automatically better.

    The inputs

    • batch_heads - size of the combined batch×heads dim (this benchmark runs a flat [BH,T,D]).
    • tokens - up to 9216, stepped by 64.
    • exact_budget - the fraction of blocks kept exact. Leave it at the default unless you know what you're doing.
    • dtype - bfloat16 only. The wheel is BF16-only; that's a hard constraint, not a choice.
    • backend - ck (default, the compiled hybrid), auto, triton, or reference, so you can isolate which layer is actually helping.

    The single benchmark STRING is the whole output. It's an output node; there's nothing to wire forward.

    Installing it - the honest version

    Pack install is the usual: Manager → search RDNA35 Attention, or git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention into custom_nodes, restart. But the ck backend needs the compiled rdna35_pisa_ck wheel under native/rdna35_pisa_ck - you build it against your exact PyTorch/ROCm runtime with MAX_JOBS=32, then pip install --no-deps it. The prebuilt wheel is narrow: Windows, BF16, gfx1151 only. If the wheel is missing, the benchmark reports the fallback instead of the hybrid. Run the pack's Diagnostics node first to confirm your stack qualifies.

    Gotchas

    That cold number is not a rounding error - the first PISA call on a fresh run can take seconds while kernels compile, and if you're benchmarking end-to-end workflow time you should either warm up or accept that the first sample is a lie. And remember the "approximate" part: even steady-state PISA output differs from dense SDPA (the pack's own generic benchmarks show cosine in the 0.70–0.81 range on random inputs; the validated 23/144 Anima profile is the quality-vetted exception). A PISA speedup is only worth anything if your model still behaves - kernel-level cosine is evidence, not a guarantee.

    CategoryRDNA35/Attention Research

    Inputs (5)

    NameTypeDefaultDescription
    batch_headsINT11–40—
    tokensINT9216128–9216—
    exact_budgetFLOAT0.15630.015625–1—
    dtypeCOMBObfloat161 options: bfloat16
    backendCOMBOck4 options: ck, auto, triton, reference

    Outputs (1)

    NameTypeDescription
    benchmarkSTRING—