Nodes/ComfyUI-sol-attn/MiniMax H3 Scheduled Sol Attention Patch
ComfyUI Node

MiniMax H3 Scheduled Sol Attention Patch

Sol-Attn on a schedule for H3

By r-vage·Created 3 days ago·Updated 3 days ago· 3
MiniMax H3 Scheduled Sol Attention Patch
  • model
  • model
  • tau_graph
◄enabledtrue►
◄tau_start1.30►
◄tau_end0.80►
◄curvelinear►
◄min_tokens4096►
◄strictfalse►
◄dense_percent0.00►
◄thresh_typediag►
◄int8_qkfalse►
◄int8_pvfalse►
◄sink_conditioningexact_kv►
◄dense_blocks►

Every sparse-attention node asks you to pick one tau and live with it for the whole render. That's a bad deal, because the two ends of a diffusion run want opposite things: early high-noise steps have loose, low-detail attention structure you can approximate freely, while the last few steps are where faces, edges and texture get placed. Sparsify those and you get mush.

This node is the MiniMax H3 Memory Efficient Sol Attention Patch with that answer baked in. Same zero-copy strided path off the fused qkv projection, same conditioning protection, same fallbacks - but tau ramps from a sparse start to a dense finish instead of sitting still. If you're running H3 and you've decided Sol is worth it, this is the one I'd wire first.

How it works

ComfyUI publishes the active timestep for each model call in transformer_options["sigmas"]. The node captures the run's sigma endpoints when it patches and maps each call onto a 0-to-1 progress value, so tau interpolates between your two ends without you configuring a step count. Go from 20 steps to 40 and the ramp stretches to match - that's the nice part.

The default direction is tau_start = 1.3 on the first, highest-noise step down to tau_end = 0.8 at the end: sparser early, denser late. curve decides how it gets there - linear (default), cosine, sqrt, smoothstep, exponential, or step, which hard-switches at the midpoint and is about as subtle as it sounds.

dense_percent is a different lever: it keeps stock dense attention for the first fraction of sampling before Sol touches anything. The Sol-Attn paper's recipe is 0.2, and the field defaults to 0 - worth trying if your early steps look unstable.

The inputs you actually touch

  • model - from your H3 loader.
  • enabled - flip to False for a clean A/B with a fixed seed.
  • tau_start (1.3), tau_end (0.8), curve - the schedule. To reproduce the non-scheduled node exactly, set both ends to the same value.
  • dense_percent (0.0) - the dense opening gate described above.
  • min_tokens (4096) - below this packed length, stock forward.
  • sink_conditioning (exact_kv) - keeps H3's packed text/conditioning/reference/audio KV blocks exact. exact_kv_and_rows also runs those query rows dense. off disables protection; if the call's layout can't be established, the node uses the dense fallback instead of guessing.
  • dense_blocks - indices kept dense, like 0-2,-1. Empty sparsifies everything.
  • strict, thresh_type, int8_qk, int8_pv - as on the other Sol nodes; strict raises on kernel errors, which is what you want for one validation run on new hardware.

Two outputs. model goes to the guider as usual. tau_graph is an IMAGE - a 512×320 plot of your schedule, with the dense gate shaded in - wire it to a Preview Image node and you can see what you actually asked for. It needs matplotlib; without it the node logs matplotlib unavailable; tau graph is blank and hands you a black frame rather than failing.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/r-vage/ComfyUI-sol-attn

Restart, or grab it through ComfyUI Manager (publisher rvage, display name "ComfyUI Sol-Attn (continued by r-vage)"). torch is the only hard dependency; Triton is required for this node and is optional only for the pack's feed-forward node. For the schedule preview:

pip install matplotlib

The pack's pyproject.toml lists matplotlib as an optional preview extra for exactly this node.

Graph order, same as its sibling: the KJNodes SageAttention patches first (global Patch Sage Attention, then MiniMax H3 Memory Efficient Sage Attention Patch), this node after them so it adopts the memory-efficient Sage forward as its fallback. Reverse that and the Sage patch shadows Sol entirely.

Where people get burned

The graph is blank. Matplotlib isn't installed. Cosmetic-only, but you lose the one view that tells you whether your ramp is sane.

It behaves like the plain H3 Sol node. It is one, if both ends are the same number. Check curve too: step is a flat line with a hard jump, which looks broken even when it isn't.

Silent dense fallbacks. With dense_percent set, early calls legitimately take the stock path and are logged as such - that's the gate doing its job. For confirmation Sol is live at all, look for [MiniMax H3 Sol] patched ... and active in the console.

The first run at a new length is slow. Triton autotunes keyed on token count, so the first pass at each new packed size pays a JIT sweep inside the loop, cached to disk afterwards. New duration, new size, pay again - benchmark the second run.

Expecting a big number. Sparsification is still approximation: you're trading some fidelity for attention throughput, and the repo's one controlled full-model pair was about 10% on steps, not a doubling. Schedule it, look at the frames, decide.

Categorymodel_patches/attention

Inputs (13)

NameTypeDefaultDescription
modelMODEL—
enabledBOOLEANtrue—
tau_startFLOAT1.300–4tau on the first, highest-noise step. Higher = more blocks take the approximate path = faster, lower fidelity.
tau_endFLOAT0.800–4tau on the final, low-noise steps where detail forms. Lower = denser attention at the end of sampling.
curveCOMBOlinearHow tau interpolates between tau_start and tau_end across sampling. step switches at the midpoint.
min_tokensINT4096256–131072Use the stock attention forward below this packed sequence length.
strictBOOLEANfalseRaise kernel errors instead of falling back. Enable while validating a new GPU or Triton version.
dense_percentFLOAT0.000–0.9Keep the stock dense attention for this fraction of early sampling (the Sol-Attn paper's recipe: 0.2). 0 disables the gate.
thresh_typeCOMBOdiagRouting threshold estimator. diag is the evaluated default; exact uses second-moment statistics for more precise routing at extra precompute cost.
int8_qkBOOLEANfalseQuantize q/k for Sol's selected exact-attention blocks. On Linux RTX 4070 Ti SUPER (SM89), measured 1.77-1.96x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.008 additional relative L2 error versus sparse Sol BF16; this excludes sparsification error versus dense attention.
int8_pvBOOLEANfalseAlso quantize the P*V dot to int8. Requires int8_qk. On Linux RTX 4070 Ti SUPER (SM89), measured 1.87-2.26x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.014 additional relative L2 error versus sparse Sol BF16. Opt-in; full-generation speed and visual quality are unmeasured.
sink_conditioningCOMBOexact_kvKeep H3's packed text/conditioning/reference/audio KV blocks exact. exact_kv_and_rows also runs those query rows dense. Missing or invalid layouts use the captured dense fallback. off disables protection. Cost depends on layout and GPU.
dense_blocksSTRINGTransformer blocks to keep dense, e.g. '0-2,-1' for the first three and the last. Negative indices count from the end. First and last blocks are the most approximation-sensitive. Empty sparsifies all.

Outputs (2)

NameTypeDescription
modelMODEL—
tau_graphIMAGE—