Patch Sol-Attn (MiniMax)
Block-sparse attention, and the three knobs that matter
- model
- MODEL
MiniMax H3 is a 33B omni-modal model that generates picture and sound in one pass. Long clips mean very long token sequences, and attention cost climbs faster than everything else you're doing. Sol-Attn is a training-free block-sparse attention route: instead of computing every query against every key, it approximates the boring part and keeps a small set of blocks exact.
The pitch is honest about when it doesn't help. Below roughly 12k tokens, dense attention is usually faster - that's the author's own note in the node description, not a hedge someone added later. The win grows with sequence length, which for H3 means longer clips, higher resolutions, or long reference payloads.
What it actually patches
The node is a MODEL patch. You put it between your model loader and your sampler, and from then on it installs an optimized_attention_override on that model plus a per-block forward hook so each transformer block knows its own index. Attention calls go through a filter first. If a call isn't eligible - not CUDA, not bf16, head_dim isn't 128, a mask is present, it's cross-attention, q and k shapes mismatch, or the sequence is shorter than min_tokens - it silently falls through to whatever attention backend you already had. That fallback is the design, not an error path: you don't get a broken graph, you get the speedup on the big calls and normal attention on the small ones.
The requirement underneath is comfy_kitchen built with the sol_attn kernel (bf16, head_dim 128, sm_80+). Missing it doesn't degrade gracefully - the node raises comfy_kitchen unavailable, or comfy_kitchen has no sol_attn; rebuild the extension (python setup.py build_ext --inplace). Worth knowing before you blame the pack.
The inputs you'll actually set
- tau (default 1.3) - the sparsity threshold. Higher is sparser: 1.0 keeps ~16% of blocks exact, 1.5 ~7%, 2.0 ~2.7%. Push it and you're trading fidelity for speed, so change one step at a time.
- min_tokens (default 12288) - leave it high. This is your guard rail against paying sparse-attention overhead on sequences where dense wins.
- sink_conditioning (default
exact_kv_and_rows) - the one that protects H3's party trick.exact_kvruns every query against the packed text/audio/reference rows exactly;exact_kv_and_rowsalso runs the target-audio query rows dense, which is what keeps generated speech and music intact. Turn it off and you're A/B testing your audio quality, probably by accident. - start_percent / end_percent (0.2 / 0.9) - dense warm-up and cool-down around the middle steps, matching the paper's recipe.
Everything else is a tuning surface you can ignore until something looks wrong. dense_blocks ("0-2,-1") pins blocks dense when you see artifacts - the first and last blocks are the most approximation-sensitive. centroid_tail defaults on for ~1.4x; flip it off for a clean quality comparison. reuse_qkv_memory writes the output into the model's fused qkv buffer and is safe because H3 discards that buffer after attention - leave it off for other models, and that caveat is real. routed_cap_percent at ~30 bounds the one workspace term that grows quadratically and is described as measured lossless.
The single output is a MODEL. Feed it to your sampler - or, if you're stacking it with the rest of this pack's routes, the pack warns about unknown combinations rather than blocking them, and the risk is yours.
Install
It ships in comfyui-minimax-h3-audio-T8, so there's no separate install:
cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8
Restart ComfyUI fully afterwards - a UI refresh won't reload the Python. The pack's requirements.txt is deliberately empty so that installing the broad node pack can't touch your Torch/CUDA stack; comfy-kitchen you install yourself. In ComfyUI Manager, search MiniMax H3 Audio T8 instead. You also need a current ComfyUI with native H3 support, since this pack leans on it.
Where people get burned
First, nothing happens and you assume it's broken. Check with verbose and look for dense ... reasons in the console - a 3-second clip at moderate resolution is genuinely below the threshold.
Second, kernel drift. Sol-Attn went through experimental comfy-kitchen builds that took max_blocks, centroid_tail and reuse_qkv_memory; the current public API doesn't. This node inspects the kernel signature and logs when it has to ignore a knob instead of failing mid-CUDA-launch. If you set routed_cap_percent and see "ignored", that's this, and it's not a bug you can fix from the workflow.
Third: don't stack attention owners. Two things claiming the same attention call is not "more speed". Watch the log - it tells you when it chains onto an existing override.
One non-technical warning: this is a niche route with essentially no community trail. The pack's name returns zero hits across the English discussion corpus, so Google won't bail you out. The repo's own docs are the documentation - and they're unusually blunt about which numbers were measured on what.
Inputs (14)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| tau | FLOAT | 1.300–4 | Threshold beta. Higher is sparser: 1.0 ~ 16% of blocks kept exact, 1.5 ~ 7%, 2.0 ~ 2.7%. |
| start_percent | FLOAT | 0.200–1 | Run dense before this point. The paper uses 0.2. |
| end_percent | FLOAT | 0.900–1 | — |
| min_tokens | INT | 122880–1048576 | Sequences shorter than this stay dense. |
| sink_conditioning | COMBO | exact_kv_and_rows | exact_kv: every query sees the packed text/audio/reference rows exactly (~3% cost). exact_kv_and_rows: additionally runs the TARGET AUDIO query rows dense (what keeps generated audio intact); reference rows stay sparse, so the cost no longer scales with reference size. |
| morton | BOOLEAN | false | Reorder video tokens into Morton (Z-order) so each 64-token block is a compact 3D neighbourhood. Exactly neutral for dense attention. |
| morton_curve | COMBO | 2d_frame | 2d_frame Z-orders within each frame and leaves frame order alone -- right for H3's non-uniform temporal axis. |
| centroid_tail | BOOLEAN | true | Evaluate the pooled branch once per 64-token query block at its centroid instead of per row. ~1.4x faster; turn OFF for a quality A/B. |
| routed_cap_percent | INT | 00–100 | Cap the routed-block list at this percent of the sequence; 0 = uncapped. Bounds the only workspace term that grows with T^2. ~30 is 3x headroom at tau=1.4 and measured lossless; below the actual density it degrades routed blocks to their pooled term. |
| reuse_qkv_memory | BOOLEAN | false | Write the attention output into the model's fused qkv buffer instead of a fresh allocation, cutting peak VRAM by one output-sized tensor (~1.2 GB at 80k tokens). Safe for MiniMax-H3, which discards that buffer after attention; leave OFF for other models unless you know theirs does too. |
| verbose | BOOLEAN | false | — |
| dense_blocks | STRING | Transformer blocks to keep dense, e.g. '0-2,-1'. The first and last blocks are the most approximation-sensitive. | |
| tau_profileopt | STRING | Per-block tau overriding the base value. 'blocks=tau' entries separated by ';' or newlines. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |