MiniMax H3 Fused Modulation
Fusing MiniMax H3's AdaLN work
- model
- MODEL
Most of the speed nodes people install for video are a negotiation: you get speed back, you pay in fidelity, and you spend an evening A/B-ing frames to find out how much. This one isn't that. MiniMax H3 Fused Modulation rewrites how H3's transformer blocks do their per-block scale/shift and gated residual arithmetic, and the output is bit-exact - the author verified torch.equal against ComfyUI's eager path on both random tensors and a real DiTBlock. Nothing to A/B. Two inputs, wire it and forget it.
How it works
A MiniMax H3 DiT block does a lot of small elementwise work per block: pull six AdaLN modulation tables (shift/scale/gate for attention, shift/scale/gate for the MLP), look up the right table row for each token segment in H3's packed layout, scale and shift the normed activation, run attention, then do a gated residual add. Repeat on the MLP side. Do that across all 50 blocks and you have a lot of kernel launches doing almost no math - each one tiny, each one paying launch overhead, each one bouncing a [tokens, hidden] tensor through memory.
This node replaces the block's forward with a fused version: four Triton launches per block instead of that stream of eager elementwise kernels. The segment-to-AdaLN-row lookup - the part that depends on the request's packed layout - is built once per layout and shared by all 50 blocks rather than recomputed per block per call.
The bit-exactness is the interesting part. ComfyUI's eager path does scale/shift in bf16 and rounds between steps, and those roundings are part of the output. So the Triton kernels reproduce the rounding explicitly, doing round-to-nearest-even on the bits rather than a cast that the compiler is free to fold away. That's why the node can claim identical output instead of "visually indistinguishable."
Fair warning on expectations: the author publishes no throughput number for this node, only that it's validated bit-exact. The win here is fewer launches and less memory traffic, and it will look small next to the attention nodes. It's the cheapest node in the pack to add, though, precisely because it can't change your image.
The inputs
model- from your H3 loader, required.enabled- your bypass.Falsereturns the model untouched.
Output is a single MODEL, straight into the guider. That's the entire interface.
It composes with everything, in either order
This node deliberately doesn't touch attention or the MLP. Those calls are resolved from the live block object at execution time, so whatever patches are installed on the attention module or the MLP - KJNodes' memory-efficient Sage, this pack's H3 Sol node, the chunked feed-forward - keep working regardless of whether fusion sits before or after them in the chain. The pack recommends it after the attention patches purely so the graph reads top to bottom.
If some other pack has replaced a block's whole forward (KJNodes' MiniMax H3 Low VRAM Attention does this), fusion leaves that block alone rather than fighting it, and logs which one it skipped.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/r-vage/ComfyUI-sol-attn
Restart ComfyUI, or install through ComfyUI Manager (publisher rvage, display name "ComfyUI Sol-Attn (continued by r-vage)"). pyproject.toml lists torch as the only hard dependency, with Triton as an optional Linux extra - but this node genuinely needs Triton, unlike the feed-forward node in the same pack, and it raises a clear error if the runtime is missing.
Requirements worth checking once: an NVIDIA GPU in the SM86/89/90/100/120/121 set, ComfyUI 0.30.1–0.37.0, and a MiniMax H3 checkpoint in ComfyUI/models/diffusion_models/.
Where people get burned
"Expected a MiniMax H3 model; returning it unchanged." You're feeding it something that isn't H3. The node checks the diffusion model's class and warns instead of doing anything clever, so a wrong wire produces a silent no-op rather than a crash.
A block silently skipped. Look for [MiniMax H3 fusion] patched 50 of 50 blocks. A lower count, or an "already has an unknown forward patch" warning, means another pack owns some block forwards and fusion stepped aside.
Eager fallback on some calls. Per-token modulation rows that are themselves tensors rather than integers need the original eager block, and so does anything that isn't a bf16 CUDA activation or that asks for gradients. The node logs [MiniMax H3 fusion] eager fallback: <reason> once per cause, and the fallback is safe - the activation is still pristine at that point.
Stacking approximations. Fusion isn't one. It changes nothing about your output, which means it does not count against the budget when you're deciding how hard to push tau on the attention nodes or how aggressive a lazy cache like EasyCache should be. Those are the dials that cost you fidelity.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| enabled | BOOLEAN | true | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |