Nodes/ComfyUI-sol-attn/MiniMax H3 Memory Efficient Sol Attention Patch
ComfyUI Node

MiniMax H3 Memory Efficient Sol Attention Patch

The H3 attention patch that doesn't copy your q/k/v three times a call

By r-vage·Created 3 days ago·Updated 3 days ago· 3
MiniMax H3 Memory Efficient Sol Attention Patch
  • model
  • MODEL
◄enabledtrue►
◄tau1.30►
◄min_tokens4096►
◄strictfalse►
◄thresh_typediag►
◄int8_qkfalse►
◄int8_pvfalse►
◄sink_conditioningexact_kv►
◄dense_blocks►

MiniMax H3 is the 33B omni-modal video model that landed in August 2026 - 4–15 second clips with native stereo audio, and heavy enough that a 16GB card runs it only in the pruned INT8 ConvRot form. Its DiT is 56 heads × 128 wide, bf16, no attention mask - precisely the shape Sol-Attn's Triton kernel wants.

It's the zero-copy sibling of the generic Sol-Attn (sparse attention) node. That one receives BHSD q/k/v through ComfyUI's attention hook and builds three contiguous copies per call - pure overhead when H3's packed sequence runs to tens of thousands of tokens. This node skips the hook: it replaces each attention module's forward and feeds the kernel strided NHD views straight off the model's fused qkv projection, reusing the same in-place RMSNorm+RoPE the stock model applies. The author's measurements put the difference at 452 MiB versus 116 MiB of peak activation per call at 8K tokens, widening to 3,612 versus 924 MiB at 65K.

How it works

Only the main DiT blocks get patched - the token refiner and short sequences behave exactly as stock. Eligible calls run the Sol kernel; ineligible ones fall back to whatever attention forward was captured when you wired the node - which is how the mem-efficient SageAttention fallback stays in play.

Then there's a problem Sol's authors never had to solve: H3 runs one joint packed sequence of text, conditioning, reference, audio and video tokens, and sparsifying the conditioning blocks would quietly break prompt adherence. That's what sink_conditioning is for.

The inputs you actually touch

  • model - wire from your H3 loader.
  • enabled - flip to False to A/B without rewiring.
  • tau (1.3) - the routing threshold. Higher is faster and looser.
  • min_tokens (4096) - below this packed sequence length the stock forward runs.
  • sink_conditioning (exact_kv) - the one H3-specific knob. exact_kv keeps the packed text/conditioning/reference/audio KV blocks exact; exact_kv_and_rows also runs those query rows dense, the most conservative option; off disables protection entirely. If the current call's packed layout can't be established, the node doesn't guess - it uses the dense fallback.
  • dense_blocks - a string of block indices to keep dense, e.g. 0-2,-1 for the first three and the last, with negatives counting from the end. Leave it empty to sparsify everything. First and last blocks are the most approximation-sensitive, so this is your lever if a render drifts.
  • strict, thresh_type, int8_qk, int8_pv - same as the generic node. strict raises instead of falling back, useful once on a new GPU; int8_pv requires int8_qk.

Output is a single MODEL. Guider next, nothing else to wire.

Order matters, and this is the part people get wrong

The pack's documented MiniMax H3 chain is:

UNETLoader → (LoRA/model patches)
           → Patch Sage Attention (KJNodes)
           → MiniMax H3 Memory Efficient Sage Attention Patch (KJNodes)
           → MiniMax H3 Memory Efficient Sol Attention Patch   ← this node
           → MiniMax H3 Fused Modulation
           → MiniMax H3 Chunk FeedForward
           → guider / sampler

Applied after the KJNodes Sage patch, this node adopts that mem-efficient Sage forward as its fallback, so ineligible steps still run mem-efficient Sage. Applied before it, the Sage patch shadows this node entirely - you get Sage everywhere plus a misleading console line. KJNodes' MiniMax H3 Low VRAM Attention is a different beast - it replaces the whole block forward, so it doesn't stack either.

This node and the Scheduled variant are alternatives, never both. The scheduled one is a superset; setting tau_start = tau_end makes them equivalent.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/r-vage/ComfyUI-sol-attn

Restart, or install through ComfyUI Manager (publisher rvage, display name "ComfyUI Sol-Attn (continued by r-vage)"). No pip install needed if ComfyUI already runs CUDA bf16 attention - torch is the only hard dependency. Triton is mandatory for this node, unlike the pack's feed-forward node: if the Sol backend failed to import, it raises rather than silently doing nothing. You'll also need an H3 checkpoint in ComfyUI/models/diffusion_models/, the pruned INT8 ConvRot one being the realistic pick.

Where people get burned

Nothing happens. Check for [MiniMax H3 Sol] patched 50 of 50 attention blocks, then active. A dense fallback: ... line names the cause: under min_tokens, conditioning protection unavailable for that call's layout, or an unsupported arch.

First-run slowness. Triton autotunes per token count, so the first pass at any new packed length pays a JIT sweep inside the sampling loop. It's cached to disk afterwards - but a new resolution or duration is a new size. Time the second run.

Wrong graph order. Covered above; it's the main way this node appears to do nothing.

Don't max everything at once. Pairing this with EasyCache is fine, but an aggressive reuse threshold plus a high tau will show up in the output. Approximations compound.

Categorymodel_patches/attention

Inputs (10)

NameTypeDefaultDescription
modelMODEL—
enabledBOOLEANtrue—
tauFLOAT1.300–4Routing threshold. Higher = more blocks take the approximate path = faster, lower fidelity. 1.0 is the Sol-Attn paper default; 1.3 is tuned here.
min_tokensINT4096256–131072Use the stock attention forward below this packed sequence length.
strictBOOLEANfalseRaise kernel errors instead of falling back. Enable while validating a new GPU or Triton version.
thresh_typeCOMBOdiagRouting threshold estimator. diag is the evaluated default; exact uses second-moment statistics for more precise routing at extra precompute cost.
int8_qkBOOLEANfalseQuantize q/k for Sol's selected exact-attention blocks. On Linux RTX 4070 Ti SUPER (SM89), measured 1.77-1.96x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.008 additional relative L2 error versus sparse Sol BF16; this excludes sparsification error versus dense attention.
int8_pvBOOLEANfalseAlso quantize the P*V dot to int8. Requires int8_qk. On Linux RTX 4070 Ti SUPER (SM89), measured 1.87-2.26x SageAttention throughput at 4K-65K tokens (tau=1, no sinks). About 0.014 additional relative L2 error versus sparse Sol BF16. Opt-in; full-generation speed and visual quality are unmeasured.
sink_conditioningCOMBOexact_kvKeep H3's packed text/conditioning/reference/audio KV blocks exact. exact_kv_and_rows also runs those query rows dense. Missing or invalid layouts use the captured dense fallback. off disables protection. Cost depends on layout and GPU.
dense_blocksSTRINGTransformer blocks to keep dense, e.g. '0-2,-1' for the first three and the last. Negative indices count from the end. First and last blocks are the most approximation-sensitive. Empty sparsifies all.

Outputs (1)

NameTypeDescription
MODELMODEL—