Nodes/Veda-on-ComfyUI/Veda Sparse Attention (MiniMax H3)
ComfyUI Node

Veda Sparse Attention (MiniMax H3)

Skip 90% of MiniMax-H3's attention and mostly get away with it

By veda-sparse·Created 3 days ago·Updated a day ago· 15
Veda Sparse Attention (MiniMax H3)
  • model
  • model
◄predictorminimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors►
◄generated_sparsity90%►
◄reference_sparsity90%►
◄full_attention_layers►
◄full_attention_steps►
◄verbosefalse►

MiniMax-H3 is a 33B omni-modal video model with native stereo audio, and the community's first question on release day was the usual one: will it run on my card. It will, slowly - and "slowly" in long-video diffusion means attention, which grows with the square of the token count. Veda computes roughly the top 10% of that attention map and skips the rest. The pack measures 2.9x end to end on an RTX 5070; a 5.2s 1344x768 clip with the 8-step Turbo LoRA goes from 342 seconds to 130.

Reach for it if you're running H3 on a 12–16GB card and your steps are dominated by waiting. Don't reach for it expecting a different model - nothing about the weights changes.

How it works

Attention is the expensive part, but most of it is wasted. A small distilled predictor - 275MB, fp8 safetensors - scores the pooled key tiles per layer, and the node keeps the top-k per query tile, builds a block mask, and hands the masked problem to a kernel that only touches the kept tiles. The kernel is Triton INT8, its arithmetic derived from SageAttention v1; the pack measured ~1.3% error against dense fp32, the same as installed SageAttention and looser than bf16 SDPA's ~0.25%. Think "the --use-sage-attention quality tier, plus sparsity", not lossless.

Two things make it more usable than a typical attention hack. It's an override through transformer_options["optimized_attention_override"], so the weights are untouched - LoRAs, quantized H3 checkpoints and the FL2VA/R2VA conditioning nodes all keep working - and it takes no patches_replace slot, so H3's Fun ControlNet still applies. And the generated video and the references get separate budgets, because one pooled top-k would starve your first frame.

The inputs that matter

Only model and predictor show until you click "show advanced inputs"; the rest defaults to the trained values, which is usually where you want to be.

  • model - the MiniMax-H3 MODEL. Put the node after any LoRA loaders and last before the guider or sampler.
  • predictor - the file from ComfyUI/models/veda/. The list is whatever .safetensors is in that folder, so a predictor trained elsewhere shows up the same way.
  • generated_sparsity (advanced, 90%) - skips 90% of the key tiles each query tile could attend. Lower is closer to full attention and slower. A whole number like 24 keeps exactly 24 key tiles of 128 tokens instead.
  • reference_sparsity (advanced, 90%) - the same knob for first/last frames, guide frames and reference images or videos. 0% keeps them in full attention, worth trying if your reference conditioning looks soft.
  • full_attention_layers / full_attention_steps (advanced, empty) - escape hatches. 0 in the steps field keeps the first sampling step dense: fix the gross structure, then let it be cheap.

There's one output, model, and it goes where the model used to. If the predictor file isn't in models/veda, the node throws instead of silently rendering dense, and the error carries the Hugging Face URL - the node never touches the network itself; the missing-models dialog is what downloads the file. After each run it tells you what it did: the kernel it picked, the size and tile plan it matched, and the share of full attention actually computed. Flip on verbose for per-phase timing and call counts.

Install

ComfyUI Manager, search "Veda", or:

cd ComfyUI/custom_nodes
git clone https://github.com/veda-sparse/Veda-on-ComfyUI

Then restart. You need ComfyUI 0.38.0 or newer - the node reads H3's packed layout out of transformer_options, so on older builds it has nothing to work with. The kernels come with the node: pyproject.toml declares triton>=3.0 (Linux) or triton-windows>=3.0 (Windows) under environment markers, plus mlx on Apple silicon. Nothing to run by hand.

If you're behind a Hugging Face mirror, set HF_ENDPOINT before restarting so ComfyUI's own download goes through it. Get the predictor either way: open Workflow → Browse Templates → Veda-on-ComfyUI, pick "Veda MiniMax H3 T2VA" and let the missing-model dialog fetch it, or drop minimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensors into ComfyUI/models/veda/ and restart.

Common issues

"Veda off: no sparse kernel works on ..." - Triton isn't installed. The node doesn't crash: it runs the model's own attention and tells you why. On Windows, Triton is the ecosystem's oldest headache and this is the failure you'll hit; pip install triton-windows (3.8.0 is the verified build) fixes it.

An error about block/head/head_dim counts - the predictor was trained against a specific H3 shape, and your checkpoint isn't it. Swap the model or the predictor; mixing them isn't supported.

Nothing changed - check you don't also have ComfyUI's own "Model Sparse Attention" node in the graph. On H3 it replaces the attention blocks outright, so Veda never gets called. The node detects that combination and warns.

You're outside the trained envelope. The predictor was trained at 1344x768, 768x1344, 768x768 and 1024x768, at 5/10/14s, with the 8-step Turbo LoRA. Other sizes fall back to the nearest trained plan - the node says nearest trained size: when they do. It still runs; compare against full attention before keeping the output.

Bypass with Ctrl+B on the same seed when you're judging output, because "looks fine" isn't "looks the same". And know the ceiling: on a 12GB card, MLP and weight movement are 69% of what's left per step, so even free attention only buys ~1.11x more.

Categorymodel/patch/minimax

Inputs (7)

NameTypeDefaultDescription
modelMODELThe MiniMax-H3 model to patch (after any LoRA loaders).
predictorCOMBOminimax_h3_t2va_veda_8nfe_600step_preview_fp8.safetensorsVeda predictor in models/veda. The official release is downloaded automatically on first use (set HF_ENDPOINT for a mirror).
generated_sparsitySTRING90%Sparsity of the generated video's attention. "90%" skips 90% of the key tiles each query tile could attend (the trained value; lower is closer to full attention and slower). A whole number such as "24" keeps exactly that many key tiles of 128 tokens instead.
reference_sparsitySTRING90%The same for the references: first / last frames, guide frames, reference images and videos. "0%" keeps full attention to and from them.
full_attention_layersSTRING0-based DiT blocks that keep full attention, e.g. "0, 1, 47-49". Empty = all sparse.
full_attention_stepsSTRING0-based sampling steps that keep full attention, e.g. "0" for the first step. Empty = all sparse.
verboseBOOLEANfalseAlso show timing and diagnostics on the node after each run, and log every decision.

Outputs (1)

NameTypeDescription
modelMODELThe model with Veda sparse attention applied.