ComfyUI Extension: ComfyUI_Monarch_Attention

Authored by xmarre

Created

Updated

3 stars

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

A custom node that enables MonarchAttention for self-attention in ComfyUI model branches by setting model-level override, affecting only the connected model branch rather than patching globally.

README

ComfyUI MonarchAttention node (model override)

This custom node enables MonarchAttention for self-attention only (where Q/K/V have the same sequence length) by setting a model-level optimized_attention_override in model_options["transformer_options"].

That means it only affects the MODEL branch you connect it to (SDXL UNet branch, WAN branch, etc.), rather than patching ComfyUI globally.

Install

  1. Copy this folder to:

    ComfyUI/custom_nodes/comfyui_monarch_attention/

    Or install via ComfyUI-Manager

Use

  • Put Enable MonarchAttention (self-attn) on the MODEL line before it goes into your sampler.
  • For multi-model workflows (e.g. WAN high/low models), place it on each model branch you want patched.
  • If something breaks, use Disable MonarchAttention (or set enable=false) on that same model branch.

Notes

  • Only applies to self-attention (square attention). It falls back to the previous attention path for cross-attention or unsupported shapes.
  • If the model already had an optimized_attention_override (e.g. SageAttention/FlashAttention), this node will chain it as a fallback and restore it when disabled.
  • Mask support is conservative (only simple boolean key-padding masks). If a call uses an unsupported mask type/shape, it will fall back.
  • impl=auto prefers triton if available; otherwise uses torch.

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Learn more