RDNA35 Patch MiniMax-H3 gfx1151 Attention
An 8x attention speedup that changes zero math — MiniMax-H3 on gfx1151, fixed by layout
- model
- model
- info
This is the node in the pack with the biggest real-world win, and the most absurdly narrow set of conditions to hit it. RDNA35 Patch MiniMax-H3 gfx1151 Attention takes a MiniMax-H3 model running on a specific AMD GPU and makes its attention calls up to ~8x faster by changing nothing about the math - no sparse approximation, no approximation at all. It just rearranges where the numbers sit in memory.
The problem it solves
MiniMax-H3 fuses its QKV projection, which means Q and K are born as interleaved views with large strides between consecutive tokens. On gfx1151 (RDNA3.5), the exact SDPA backend is brutally slow when it has to repeatedly walk those strides. The fix is embarrassingly simple in hindsight: pack KV (for 4608 <= T < 12000) or QKV (for 12000 <= T <= 16500) into one contiguous torch.stack allocation, then call the existing attention backend with the original arguments. That's the whole mechanism. No kernel replacement, no approximating, no touching the model's weights - the README's measured outputs are bit-identical (max abs error 0, cosine 1.0), and the correctness gate is a BF16 allclose at 5e-2.
The inputs and outputs
Three inputs, all boring:
model- the MiniMax-H3 MODEL, wired in after your attention-backend/loader nodes.enabled- true by default; flip it to bypass the patch.verbose_fallbacks- log fallback reasons to console when on.
Outputs are model (the cloned, patched model - wire it onward) and info (a STRING stating whether the H3 patch installed, and if not, why). If a container-aware model-local override is already installed, the node refuses to install rather than bypass that contract, and tells you so in info. The original model is returned untouched on any failure. That safety posture is the pack's signature: model-local clone, never a global monkey-patch, fallbacks chained.
Where the conditions bite
The README is brutally specific about when this node is allowed to act: ROCm gfx1151, BF16, B=1, H=56, D=128, merged output, no mask, no GQA, no causal mode, no gradients, forward-only, 4608 <= T <= 16500, and only native comfy.ldm.minimax.model.MiniMaxH3Model calls. Below 4608 or above 16500 tokens - where packing isn't validated as beneficial - it leaves the call on the existing path. T2V's integrated MiniMax-H3 node doesn't even expose a MODEL socket, so this patch only plugs into R2V workflows with a visible model chain (the README ships one). If your model selects FP16 attention, the BF16-only condition fails and you get a clean fallback, not a crash.
Installing it
Manager → search RDNA35 Attention, or:
cd ComfyUI/custom_nodes
git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention
Restart. Base install is zero-dependency. The measured results assume a PyTorch 2.14 / ROCm 7.15 stack with AOTriton - the changelog notes it deliberately picks SDPA/AOTriton over the CUDA-only Comfy Kitchen SageAttention path for this workflow. The actual VRAM cost is real, though: packing alone allocates ~251 MiB at T=9170 and ~677 MiB at T=16500 on top of everything else, so don't run it half a gigabyte from an OOM.
Honest take
If you run MiniMax-H3 on an RX 9070-series card, this is the node worth your time - it's a rare "free" speedup with bit-identical output, which is almost unheard of in the attention world. If you don't have gfx1151, or you don't run MiniMax-H3, the node is a no-op on your hardware and no amount of configuration will change that. Check the architecture with the pack's Diagnostics node first; on anything else, this patch quietly does nothing while remaining entirely safe.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| enabled | BOOLEAN | true | — |
| verbose_fallbacks | BOOLEAN | false | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| info | STRING | — |