Nodes/ComfyUI RDNA35 Attention/RDNA35 Fixed Block Attention Benchmark
ComfyUI Node

RDNA35 Fixed Block Attention Benchmark

Benchmarking block-diagonal attention without burning a single real image

By Yasei-no-otoko·Created 3 months ago·Updated 2 months ago· 1
RDNA35 Fixed Block Attention Benchmark
    • benchmark
    ◄batch1►
    ◄heads4►
    ◄tokens128►
    ◄head_dim64►
    ◄dtypefloat16►
    ◄modeauto►
    ◄causalfalse►

    If you want to know whether fixed 64-token block-diagonal attention is worth anything on your GPU, you do not want to find out by running a 30-step generation and squinting at two images. You want a controlled, synthetic benchmark where everything else is held constant and the only variable is the attention path. That's exactly what RDNA35 Fixed Block Attention Benchmark is: it fabricates random Q/K/V tensors, runs four attention implementations on them, and hands you a text report with medians, errors, and the honest semantic caveats.

    What it measures

    For your chosen shape it times:

    • the PyTorch reference loop for fixed block-diagonal attention (the correctness baseline),
    • the dispatch path - Triton when the pack's conditions hold (ROCm build, Triton importable, fp16/bf16, head dim 32/64/128, matching self-attention shapes, no mask, no gradients), otherwise the reference,
    • PyTorch SDPA with a block-diagonal mask, which has the same fixed-block semantics and is your fair "optimized conventional" comparison,
    • normal full SDPA - included as a semantic contrast, not a competitor, because block-diagonal and full attention compute different things.

    That last one is the trap the node is actively protecting you from. Full attention is faster and different - it keeps cross-block information. The report prints the "full SDPA semantic delta vs fixed-block reference" so you can't mistake a speed ratio for a like-for-like comparison.

    The inputs that matter

    Only a few, and they map straight onto the shape of the fake tensors:

    • tokens - sequence length, stepped by 64 (block-aligned), up to 8192.
    • head_dim - 32, 64, or 128; the Triton kernel only supports these.
    • dtype - float16 or bfloat16.
    • mode - auto (let dispatch decide), triton (force it), reference (skip the kernel entirely).
    • batch, heads, and causal round it out.

    The single benchmark STRING output prints device, torch/HIP versions, Triton status, the reference and SDPA medians, the dispatch backend, the fallback reason if any, and the max abs error against the reference. If the Triton path didn't run, it says so outright: "measured speedup: unavailable because optimized path did not run." No hiding.

    Installing and running it

    Manager → search RDNA35 Attention, or clone the repo into custom_nodes:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention
    

    Restart, and it lives under RDNA35/Fixed Block Attention. Zero pip dependencies in the base install. It also runs on CPU - the reference path is pure PyTorch - which makes it a great sanity check on a machine that can't run the pack's optimized kernels at all.

    Gotchas worth knowing

    First, this is a kernel-level diagnostic, not an end-to-end image test. A Triton win here tells you nothing about whether block-sparse attention degrades your actual outputs - that's a model-level question this node deliberately doesn't answer. Second, if RDNA35 Block Attention Diagnostics says your machine isn't RDNA3.5 with working Triton, expect the "unavailable" line; the node is still useful as a baseline collector. And third, this pack is a single-author research project with a small footprint - the README's own tables are the closest thing to reference results, and they're measured on the author's gfx1151 Windows ROCm stack. Treat your numbers as machine-specific, which is the point of having a benchmark at all.

    CategoryRDNA35/Fixed Block Attention

    Inputs (7)

    NameTypeDefaultDescription
    batchINT11–8—
    headsINT41–32—
    tokensINT1281–8192—
    head_dimINT6432–128—
    dtypeCOMBOfloat162 options: float16, bfloat16
    modeCOMBOauto3 options: auto, triton, reference
    causalBOOLEANfalse—

    Outputs (1)

    NameTypeDescription
    benchmarkSTRING—