Nodes/H3 Optimizations/H3 VSA Attention (FastH3)
ComfyUI Node

H3 VSA Attention (FastH3)

A FastH3 checkpoint already learned its sparsity — this node just lets it use it

By Zironic·Created 2 months ago·Updated 11 days ago· 127
H3 VSA Attention (FastH3)
  • model
  • MODEL
◄keep_percent20.0►
◄dense_first_steps0►
◄dense_layers►
◄backendAuto►
◄memory_modeStandard►
◄verbosefalse►

MiniMax H3's 33B video model is expensive to sample at full length, so people chop up the attention. The pack's main "H3 Sparse Attention" node does that on a knob: pick a Video attention budget and it drops video-to-video attention connections by percentage. Fine for the base checkpoint.

The FastH3 checkpoints weren't retrofitted with sparsity - they were trained with it, gates and all. Point the generic sparse node at one and it detects the learned gate layers, logs that the checkpoint "was trained on its own sparse pattern," and passes the model through untouched. Zero speedup, one warning line. H3 VSA Attention (FastH3) is the node that actually runs those checkpoints.

What it does, mechanically

VSA is video sparse attention, and this node reproduces FastVideo's VSA-H3 geometry exactly rather than offering its own cube orders or schedules. The target video is tiled into 4×4×4 (t, h, w) cubes of 64 rows; text and audio get their own 64-row prefix tiles so no tile straddles two modalities.

Per block, Q/K/V are projected in bounded chunks, RMS-normed, RoPE'd, and scattered into those tile buffers - the full fused Q/K/V temporary never exists, which is half the point. fp32 means of post-RoPE Q, K and V per tile then give tile scores of q_mean · k_mean / sqrt(d). Prefix query tiles attend everything (text and audio stay dense, always). Video query tiles attend every prefix tile plus their top-k video tiles, k = ceil((1 - sparsity) * n).

The clever bit is the coarse branch: whatever top-k threw away is summarized by softmax(scores) @ v_mean over all tiles, added to every row and scaled per token by the checkpoint's own to_gate_compress gate. The model blends a cheap global summary with its sparse local view instead of just losing the dropped context.

The inputs that matter

One field decides whether this works: Video tiles kept (%), default 20, range 1–100. Set it to what the checkpoint was trained at - the author's tooltip names 20 for FastH3 8-Step V2 (80% sparse) and 10 for FastH3 Preview v1 (90% sparse). Freehanding it moves you off the trained operating point; higher is slower and not guaranteed to look better.

The rest is diagnostics unless you're debugging:

  • dense_first_steps (default 0) and dense_layers (default empty, format 0, 1, 47-49) - sampler steps, or DiT blocks, that keep every video tile. FastH3 was trained sparse on every step, so the defaults match training.
  • backend - Auto uses the vendored Kitchen INT8 tile kernel when it's present and passes a one-time numerical check against BF16 on your GPU, otherwise BF16 Triton.
  • memory_mode - Standard vs Lower VRAM (slower), INT8-only: it stages V in two passes and writes attention output into the block input it replaces. For when long sequences don't fit.
  • verbose - logs the tile geometry once per shape.

Output is a single MODEL, wired into whatever feeds the sampler. Sit H3 Memory Optimization next to it - the two compose, because the VSA node hands ComfyUI's attention slot a VSA callable over the memory node's block forward. Order barely matters: the pack keeps its plan on the cloned ModelPatcher, so a LoRA loader downstream won't quietly break it.

Install

ComfyUI Manager, search H3 Optimizations, or manually:

cd ComfyUI/custom_nodes
git clone https://github.com/Zironic/H3-Optimizations

Restart and the nodes appear under H3-Optimizations > Model Patches. There's nothing to pip install - no declared dependencies, and the INT8 kernels ship as prebuilt binaries in native/bin; nothing downloads or compiles at startup. You need ComfyUI 0.33.0 or newer and Python 3.10+.

The catch is what the pack doesn't ship: this is a model patch, not a loader. The checkpoint is on you.

Where people get burned

Wrong node, two directions. Point this at a stock H3 checkpoint and it raises immediately, telling you to use H3 Sparse Attention instead. Point H3 Sparse Attention at a FastH3 checkpoint and it warns and passes through, which reads as "sparse attention isn't doing anything." Both are the pack being honest.

Auto backend isn't slow - it fell back. The INT8 path needs the shipped native library to load cleanly and pass a per-GPU parity check; a missing or mismatched build warns and quietly uses BF16 Triton. On ROCm you're on BF16 Triton period, because the AMD library doesn't expose the VSA tile entry point. The CUDA targets cover sm75, sm80, sm89, sm120 and Linux sm90a plus a compute_89 PTX fallback, and only sm89 has had a live GPU run for this kernel.

A dense prefix on a distilled checkpoint. The name says 8-step; like every distilled video model it wants the step count and schedule it was trained on, and the tooltip is explicit that FastH3 was trained sparse on every step. dense_first_steps above 0 is a diagnostic override, not a free quality bump.

Version reality. The README says ComfyUI 0.33.0+, but the source demands H3 blocks whose forward accepts an attention replacement - 0.35 or newer. If you get that error rather than the checkpoint error, update ComfyUI.

Last thing before you spend GPU hours on H3: the model's community licence excludes the US, EU, UK and South Korea from its applicable territory. The node doesn't care. Your lawyer might.

CategoryH3-Optimizations/Model Patches

Inputs (7)

NameTypeDefaultDescription
modelMODEL—
keep_percentFLOAT20.01–100Percent of video tiles each video query tile attends exactly. Use the value the checkpoint was trained at: 20 for FastH3 8-Step V2 (80% sparse), 10 for FastH3 Preview v1 (90% sparse). Other values leave the trained operating point; more is slower and not guaranteed to look better.
dense_first_stepsINT00–1000Sampler steps that keep every video tile. FastH3 was trained sparse on every step, so 0 matches training.
dense_layersSTRINGBlocks that keep every video tile, e.g. '0, 1, 47-49'. Empty matches training.
backendCOMBOAutoKernel for the exact sparse pass. Auto uses the vendored Kitchen INT8 kernel when it is present and passes a one-time check against BF16 on this GPU, else BF16 Triton. Tile selection and the coarse branch always run in fp32.
memory_modeCOMBOStandardINT8 only. Lower VRAM stages V in two passes instead of keeping it in BF16, and writes attention output into the block input it replaces. Slower; use it when the long sequences do not fit.
verboseBOOLEANfalseLog the tile geometry once per shape.

Outputs (1)

NameTypeDescription
MODELMODEL—