Nodes/ComfyUI RDNA35 Attention/RDNA35 Patch Model Attention
ComfyUI Node

RDNA35 Patch Model Attention

Block-diagonal attention, the safe way — this node patches your model without breaking everything else

By Yasei-no-otoko·Created 3 months ago·Updated 2 months ago· 1
RDNA35 Patch Model Attention
  • model
  • model
  • info
◄enabledtrue►
◄modeauto►
◄semantic_modeexact_only►
◄block_size64►
◄causalfalse►
◄verbose_fallbacksfalse►

Every other attention speedup pack in the ComfyUI ecosystem monkey-patches optimized_attention globally and hopes for the best. This one refuses. RDNA35 Patch Model Attention installs a model-local attention override on a cloned copy of your MODEL, and by default it does absolutely nothing to any call that doesn't explicitly ask for it. That's not a bug. It's the entire design.

The idea it implements is fixed 64-token block-diagonal self-attention: tokens in block i only attend to keys and values from block i, so each 64-token window is attended to in isolation. It's the same research idea as PyTorch's TLX block attention blog post the README links, and it's a real algorithmic change - cross-block attention is removed, so the output is not equivalent to normal full attention. You're trading information flow for compute. The name is honest about that, which is rare.

How it works

You feed it a MODEL, it clones the model, and it installs an optimized_attention_override that only converts calls explicitly marked as fixed_block_diagonal - via a rdna35_attention_semantics flag or the matching transformer_options entry. Everything else passes through untouched to the previous backend. Cross-attention is never converted, period.

The two modes:

  • exact_only (default): only explicitly marked calls get converted. Ordinary ComfyUI attention is untouched. This is the safe mode and the reason the node won't silently change your images.
  • experimental_force_block_local: opt-in, and it still refuses to convert anything it can't prove is self-attention (it checks a flag, then falls back to pointer-identity - q.data_ptr() == k.data_ptr()).

The other inputs: mode (auto picks Triton when the dispatch conditions hold, otherwise the reference implementation - or you can force triton/reference), causal for causal masks, block_size which is locked to 64 (no point wiggling it), and verbose_fallbacks to log the reasons to console.

The outputs that matter

  • model - the cloned, patched model. Wire this into your sampler instead of the original. If a safe patch can't be installed, the node returns the original model untouched.
  • info - a STRING telling you what actually happened: whether the override was installed, or why it wasn't (e.g. a container-aware override already present). Check this before you blame your sampler for a non-speedup.

Installing it

Manager, search RDNA35 Attention, or:

cd ComfyUI/custom_nodes
git clone https://github.com/Yasei-no-otoko/ComfyUI-RDNA35-Attention

Restart, and it appears under RDNA35/Fixed Block Attention. The base install has zero pip dependencies. The Triton path needs a PyTorch ROCm build, a matching Triton, and an RDNA3.5 GPU (gfx1150/1151/1152) - run the pack's Diagnostics node first to check. Without those, mode: triton silently falls back to a PyTorch reference implementation, which is slower than the ComfyUI backend it replaced.

Where people get burned

The most common mistake is expecting a speedup by wiring this in with exact_only and normal workflows. You won't get one, because no stock ComfyUI call marks itself fixed_block_diagonal - the node deliberately keeps you from converting anything by accident. To actually use it you have to opt a call into the semantics, which means this is research plumbing, not a drop-in accelerator. That's the honest framing: a safe, model-local implementation of a research attention pattern, built for people experimenting with block-sparse attention on AMD RDNA3.5, not a "make everything faster" button. If that's what you wanted, keep scrolling.

CategoryRDNA35/Fixed Block Attention

Inputs (7)

NameTypeDefaultDescription
modelMODEL—
enabledBOOLEANtrue—
modeCOMBOauto3 options: auto, reference, triton
semantic_modeCOMBOexact_only2 options: exact_only, experimental_force_block_local
block_sizeINT6464–64—
causalBOOLEANfalse—
verbose_fallbacksBOOLEANfalse—

Outputs (2)

NameTypeDescription
modelMODEL—
infoSTRING—