Nodes/ComfyUI-sol-attn/Sol-Attn (sparse attention)
ComfyUI Node

Sol-Attn (sparse attention)

The sparse-attention patch you drop between your loader and your guider

By r-vage·Created 3 days ago·Updated 3 days ago· 3
Sol-Attn (sparse attention)
  • model
  • MODEL
◄enabledtrue►
◄tau1.30►
◄min_tokens4096►
◄strictfalse►
◄thresh_typediag►
◄int8_qkfalse►
◄int8_pvfalse►

Sol-Attn doesn't generate anything, and it isn't a sampler. It's a MODEL in, MODEL out patch: you wire it between your checkpoint loader and your guider and it quietly reroutes the model's self-attention through NVIDIA's Sol-Attn Triton kernel. Bypass it and your graph is stock again. That's the whole interface - and the reason to like it.

One thing to get straight first: on MiniMax H3 this is the wrong node from the pack. Use MiniMax H3 Memory Efficient Sol Attention Patch or its Scheduled variant. This generic one is for everything else - Wan, LTX, Flux - and on H3 it leaves real performance and memory on the table.

How it works

The node writes optimized_attention_override into the model's transformer_options. Every attention call in that model then hits Sol's override first, where it acts as a gatekeeper before it acts as a speedup. The kernel needs head_dim of exactly 128, bf16 q/k/v on one CUDA device, no attention mask, the 4D skip_reshape call path, an arch in SM86/89/90/100/120/121, and at least min_tokens tokens. Anything else gets handed to whatever backend was already there - your SageAttention, your ComfyUI cross-attention setting. Each distinct failure reason is logged once per run, not once per call.

The kernel itself is content-dependent sparsification: K and V get summarized into 64-token blocks, the blocks get scored against each query, and the ones that don't matter take an approximate path. tau is where you draw that line - standard deviations above the mean block score. Higher tau means more blocks go cheap: faster, looser. The author's default here is 1.3; the Sol-Attn paper's is 1.0.

One honest detail: Comfy hands the hook BHSD tensors and the kernel wants contiguous BTHD, so the node copies q, k and v on every call - 0.2–2 ms depending on length. That's why the H3 node in this pack exists.

The inputs you actually touch

  • model - from your loader. Required.
  • enabled - your A/B switch. Keep it wired and flip to False to compare against dense with a fixed seed.
  • tau (1.3) and min_tokens (4096) - the two real knobs. Below min_tokens the node does nothing at all.
  • strict - off by default. Turn it on once on new hardware: instead of silently falling back it raises, which is how you find out your GPU or Triton version genuinely isn't supported.
  • thresh_type - diag (the evaluated default) or exact, which uses second-moment statistics for sharper routing at extra precompute cost.
  • int8_qk, int8_pv - the quantized paths. int8_pv requires int8_qk.

The only output is a MODEL. It goes into your guider, nothing else.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/r-vage/ComfyUI-sol-attn

Restart ComfyUI, or search ComfyUI Manager for the pack - pyproject.toml registers it under publisher rvage as "ComfyUI Sol-Attn (continued by r-vage)". No requirements.txt, no pip step: the only hard dependency is torch, which ComfyUI already supplies, with Triton as an optional Linux extra. On Windows, use woct0rdho's Triton builds rather than a fresh Linux wheel.

You need an NVIDIA card in the SM86/89/90/100/120/121 range, PyTorch with CUDA and bf16, and ComfyUI 0.30.1–0.37.0. The pack is a maintained continuation of the original Saganaki22 (drbaph) integration whose repo is gone; keep one copy installed, since it preserves the original node IDs.

Where people get burned

"I installed it and nothing changed." Read the console. You want [Sol-Attn] patched (tau=1.30, ...) at load time, then [Sol-Attn] active during sampling. If you see [Sol-Attn] dense fallback: <reason>, the reason is printed right there - usually a sequence under min_tokens, a head dim that isn't 128, or an attention mask on the call.

The first run at a new size is slow, and that isn't the node misbehaving. Triton autotunes keyed on token count, so the first sampling pass at any new token count pays a JIT sweep inside the loop. It's cached to disk, so you pay it once per size - but a new duration or resolution is a new size. Benchmark the second run.

Don't expect the headline multiplier end to end. The published 1.48–1.56× is attention throughput against SageAttention 2.2.0 on one RTX 4070 Ti SUPER. The same repo's single controlled full-model pair went 9.91 s/it → 8.92 s/it, about 10%. Below roughly 4K tokens SageAttention wins outright.

Stacking it with the H3 nodes is wrong. KJNodes' H3 memory-efficient Sage patch replaces each attention module's forward directly, bypassing the global hook this node uses - so the H3 Sol node is the one to stack there.

It's approximate, and more than you'd guess. The author's own measurement puts the sparse path at 8K tokens at 0.750 relative L2 against dense SDPA - a real change in output, not rounding noise. A/B it with enabled=False first.

Categorymodel_patches/attention

Inputs (8)

NameTypeDefaultDescription
modelMODEL—
enabledBOOLEANtrue—
tauFLOAT1.300–4Routing threshold. Higher = more blocks take the approximate path = faster, lower fidelity. 1.0 is the Sol-Attn paper default; 1.3 is tuned here.
min_tokensINT4096256–131072Use the normal backend below this sequence length. Linux RTX 4070 Ti SUPER (SM89) benchmarks cover 4K-65K tokens.
strictBOOLEANfalseRaise kernel errors instead of falling back. Enable while validating a new GPU or Triton version.
thresh_typeCOMBOdiagRouting threshold estimator. diag is the evaluated default; exact uses second-moment statistics for more precise routing at extra precompute cost.
int8_qkBOOLEANfalseQuantize q/k for Sol's selected exact-attention blocks. On Linux RTX 4070 Ti SUPER (SM89), measured 1.77-1.96x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.008 additional relative L2 error versus sparse Sol BF16; this excludes sparsification error versus dense attention.
int8_pvBOOLEANfalseAlso quantize the P*V dot to int8. Requires int8_qk. On Linux RTX 4070 Ti SUPER (SM89), measured 1.87-2.26x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.014 additional relative L2 error versus sparse Sol BF16. Opt-in; full-generation speed and visual quality are unmeasured.

Outputs (1)

NameTypeDescription
MODELMODEL—