MiniMax H3 Scheduled Sol Attention Patch
Sol-Attn on a schedule for H3
- model
- model
- tau_graph
Every sparse-attention node asks you to pick one tau and live with it for the whole render. That's a bad deal, because the two ends of a diffusion run want opposite things: early high-noise steps have loose, low-detail attention structure you can approximate freely, while the last few steps are where faces, edges and texture get placed. Sparsify those and you get mush.
This node is the MiniMax H3 Memory Efficient Sol Attention Patch with that answer baked in. Same zero-copy strided path off the fused qkv projection, same conditioning protection, same fallbacks - but tau ramps from a sparse start to a dense finish instead of sitting still. If you're running H3 and you've decided Sol is worth it, this is the one I'd wire first.
How it works
ComfyUI publishes the active timestep for each model call in transformer_options["sigmas"]. The node captures the run's sigma endpoints when it patches and maps each call onto a 0-to-1 progress value, so tau interpolates between your two ends without you configuring a step count. Go from 20 steps to 40 and the ramp stretches to match - that's the nice part.
The default direction is tau_start = 1.3 on the first, highest-noise step down to tau_end = 0.8 at the end: sparser early, denser late. curve decides how it gets there - linear (default), cosine, sqrt, smoothstep, exponential, or step, which hard-switches at the midpoint and is about as subtle as it sounds.
dense_percent is a different lever: it keeps stock dense attention for the first fraction of sampling before Sol touches anything. The Sol-Attn paper's recipe is 0.2, and the field defaults to 0 - worth trying if your early steps look unstable.
The inputs you actually touch
model- from your H3 loader.enabled- flip toFalsefor a clean A/B with a fixed seed.tau_start(1.3),tau_end(0.8),curve- the schedule. To reproduce the non-scheduled node exactly, set both ends to the same value.dense_percent(0.0) - the dense opening gate described above.min_tokens(4096) - below this packed length, stock forward.sink_conditioning(exact_kv) - keeps H3's packed text/conditioning/reference/audio KV blocks exact.exact_kv_and_rowsalso runs those query rows dense.offdisables protection; if the call's layout can't be established, the node uses the dense fallback instead of guessing.dense_blocks- indices kept dense, like0-2,-1. Empty sparsifies everything.strict,thresh_type,int8_qk,int8_pv- as on the other Sol nodes;strictraises on kernel errors, which is what you want for one validation run on new hardware.
Two outputs. model goes to the guider as usual. tau_graph is an IMAGE - a 512×320 plot of your schedule, with the dense gate shaded in - wire it to a Preview Image node and you can see what you actually asked for. It needs matplotlib; without it the node logs matplotlib unavailable; tau graph is blank and hands you a black frame rather than failing.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/r-vage/ComfyUI-sol-attn
Restart, or grab it through ComfyUI Manager (publisher rvage, display name "ComfyUI Sol-Attn (continued by r-vage)"). torch is the only hard dependency; Triton is required for this node and is optional only for the pack's feed-forward node. For the schedule preview:
pip install matplotlib
The pack's pyproject.toml lists matplotlib as an optional preview extra for exactly this node.
Graph order, same as its sibling: the KJNodes SageAttention patches first (global Patch Sage Attention, then MiniMax H3 Memory Efficient Sage Attention Patch), this node after them so it adopts the memory-efficient Sage forward as its fallback. Reverse that and the Sage patch shadows Sol entirely.
Where people get burned
The graph is blank. Matplotlib isn't installed. Cosmetic-only, but you lose the one view that tells you whether your ramp is sane.
It behaves like the plain H3 Sol node. It is one, if both ends are the same number. Check curve too: step is a flat line with a hard jump, which looks broken even when it isn't.
Silent dense fallbacks. With dense_percent set, early calls legitimately take the stock path and are logged as such - that's the gate doing its job. For confirmation Sol is live at all, look for [MiniMax H3 Sol] patched ... and active in the console.
The first run at a new length is slow. Triton autotunes keyed on token count, so the first pass at each new packed size pays a JIT sweep inside the loop, cached to disk afterwards. New duration, new size, pay again - benchmark the second run.
Expecting a big number. Sparsification is still approximation: you're trading some fidelity for attention throughput, and the repo's one controlled full-model pair was about 10% on steps, not a doubling. Schedule it, look at the frames, decide.
Inputs (13)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| enabled | BOOLEAN | true | — |
| tau_start | FLOAT | 1.300–4 | tau on the first, highest-noise step. Higher = more blocks take the approximate path = faster, lower fidelity. |
| tau_end | FLOAT | 0.800–4 | tau on the final, low-noise steps where detail forms. Lower = denser attention at the end of sampling. |
| curve | COMBO | linear | How tau interpolates between tau_start and tau_end across sampling. step switches at the midpoint. |
| min_tokens | INT | 4096256–131072 | Use the stock attention forward below this packed sequence length. |
| strict | BOOLEAN | false | Raise kernel errors instead of falling back. Enable while validating a new GPU or Triton version. |
| dense_percent | FLOAT | 0.000–0.9 | Keep the stock dense attention for this fraction of early sampling (the Sol-Attn paper's recipe: 0.2). 0 disables the gate. |
| thresh_type | COMBO | diag | Routing threshold estimator. diag is the evaluated default; exact uses second-moment statistics for more precise routing at extra precompute cost. |
| int8_qk | BOOLEAN | false | Quantize q/k for Sol's selected exact-attention blocks. On Linux RTX 4070 Ti SUPER (SM89), measured 1.77-1.96x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.008 additional relative L2 error versus sparse Sol BF16; this excludes sparsification error versus dense attention. |
| int8_pv | BOOLEAN | false | Also quantize the P*V dot to int8. Requires int8_qk. On Linux RTX 4070 Ti SUPER (SM89), measured 1.87-2.26x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.014 additional relative L2 error versus sparse Sol BF16. Opt-in; full-generation speed and visual quality are unmeasured. |
| sink_conditioning | COMBO | exact_kv | Keep H3's packed text/conditioning/reference/audio KV blocks exact. exact_kv_and_rows also runs those query rows dense. Missing or invalid layouts use the captured dense fallback. off disables protection. Cost depends on layout and GPU. |
| dense_blocks | STRING | Transformer blocks to keep dense, e.g. '0-2,-1' for the first three and the last. Negative indices count from the end. First and last blocks are the most approximation-sensitive. Empty sparsifies all. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |
| tau_graph | IMAGE | — |