Nodes/ComfyUI RDNA35 Attention/RDNA35 MiniMax-H3 gfx1151 Attention Benchmark
ComfyUI Node

RDNA35 MiniMax-H3 gfx1151 Attention Benchmark

A benchmark that refuses to run on your GPU — and that's the point

By Yasei-no-otoko·Created 3 months ago·Updated 2 months ago· 1
RDNA35 MiniMax-H3 gfx1151 Attention Benchmark
    • benchmark
    ◄tokens9170►

    Here's a node that will just refuse to run on 99% of the GPUs out there. RDNA35 MiniMax-H3 gfx1151 Attention Benchmark checks your GPU's architecture, and if it isn't gfx1151, it raises an error before doing anything. That's not broken behavior - it's a hard requirement, because this benchmark measures something that only exists on one GPU.

    MiniMax-H3 is MiniMax's open-weights video model, and its attention has a quirk: Q and K come out of a fused QKV projection as interleaved views, with large strides between tokens. On RDNA3.5's gfx1151 silicon, the exact SDPA backend is pathologically slow at repeatedly walking those strides. The pack's patch node (RDNA35 Patch MiniMax-H3 gfx1151 Attention) repacks the tensors into a contiguous layout before calling the backend, and this benchmark exists to quantify the win: native interleaved layout vs. packed layout, same inputs, same SDPA backend, measured head-to-head.

    What it does

    One input: tokens, from 4608 to 16500. It builds a synthetic B=1, H=56, D=128 BF16 QKV tensor - the exact shape the patch targets - and times both the native interleaved path and the packed path with alternating GPU events (2 warmups, 5 samples; the README notes the CLI script with 20 samples is the canonical measurement, this is the quick in-UI diagnostic). It runs the pack's pack_minimax_h3_qkv logic, which packs KV for 4608 <= T < 12000 and QKV for 12000 <= T <= 16500.

    Crucially, it doesn't trust its own timings. Before timing anything, it checks that the packed output is finite and passes a BF16 correctness gate (allclose with atol/rtol 5e-2) against the native path, and raises if not. The single benchmark STRING output then reports both medians, the speedup, the time reduction, max abs error, cosine similarity, and which packing profile applied. On the author's stack (PyTorch 2.14 / ROCm 7.15, AOTriton, gfx1151) those numbers are 3.6–8x with bit-identical output - at T=16500, 2264 ms dropping to 280 ms.

    Installing and running it

    ComfyUI Manager → search RDNA35 Attention, or:

    cd ComfyUI/custom_nodes
    git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention
    

    Restart, find it under RDNA35/Attention Research. Zero base dependencies. The runtime needs are the whole story: a PyTorch ROCm build with torch.version.hip set, an actual gfx1151 GPU, and a working Triton/AOTriton for the SDPA backend. On anything else the error message tells you exactly what was missing.

    Gotchas

    Don't install this pack for this benchmark alone - the benchmark is evidence for the patch node, which is where the actual speedup happens in a real MiniMax-H3 workflow. And note what the benchmark does not measure: the timed region includes packing plus the SDPA call, but the packing alone allocates ~251 MiB at T=9170 and ~677 MiB at T=16500 on top of model tensors and backend workspace. That's real VRAM pressure, and a VRAM-limited card could choke on the QKV-packed sizes even though the kernel math is identical. Read the numbers as "attention call got faster," not "your whole video got 8x faster."

    CategoryRDNA35/Attention Research

    Inputs (1)

    NameTypeDefaultDescription
    tokensINT91704608–16500—

    Outputs (1)

    NameTypeDescription
    benchmarkSTRING—