Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Low VRAM Attention / 低显存注意力 (Advanced EXP/T8)
ComfyUI Node

MiniMax H3 Low VRAM Attention / 低显存注意力 (Advanced EXP/T8)

Head-grouping to shave H3's attention peak

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Low VRAM Attention / 低显存注意力 (Advanced EXP/T8)
  • model
  • model
  • report_json
◄head_chunks4►

Attention is where a video diffusion model's peak memory actually lives, and H3 generates video and audio jointly, so H3's attention is doing more work than most. This node attacks that peak from two directions: it releases the normalized block input earlier than stock H3 does, and it splits the self-attention call across head groups so no single kernel launch allocates the whole temporary at once. The sibling of this node splits the FFN; together they're the pack's "my 16GB card is the wall" pair.

The mechanism

Both changes happen inside a clone of the MODEL you pass in, not globally. The formula is untouched - this is the same attention H3 already computes - but the launch shape changes, so floating-point rounding changes with it. Two runs at the same seed with this node on will not be pixel-identical to a run with it off. Do not A/B this thing with a pixel diff; look at the frames.

Worth understanding what it doesn't do: it never imports KJNodes and never alters ComfyUI's process-global attention choice. If you launch with --use-sage-attention or enable a backend per workflow, the node keeps your existing override - it calls it once per head group and reports what it did. That's a real guarantee that your backend isn't silently dropped, and it is not a guarantee that every third-party backend behaves well when handed grouped heads. Sol and Sage need a short-clip check on your exact version.

Inputs and outputs

The whole node is two fields and two outputs.

  • model - the H3 model, after any ordinary weight LoRA.

  • head_chunks (default 4, 1–56) - how many groups to split the attention heads into. Bigger generally means a lower per-call temporary and more launch overhead; the author's own framing is "not a case of bigger is better." At 1 you still get the early input release, you just lose the grouping. The effective group count is capped by the model's actual head count, so the 56 ceiling is only theoretical.

  • model out goes to your sampler, or into the Chunk FeedForward node - either order works, and either node works alone.

  • report_json is a string describing what was attached. Handy when you're trying to tell whether your backend actually stayed in the chain.

Install

Manager, search MiniMax H3 Audio T8, install, restart completely. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

No pip packages - the pack's requirements file is empty on purpose so it can't swap out ComfyUI's torch. This node does require a recent ComfyUI core: it validates the DiTBlock, Attention and MLP forward signatures at load and bails with a clear error if your core is the wrong shape. T8mars is the author behind comfyui-purgevram, the VRAM-clearing pack people recommend when someone asks how to force ComfyUI to unload models, so low-memory plumbing is their lane.

The honest expectation

Paired with ChunkFeedForward(chunks=2), the author measured a fixed 3-second 1024×512, 73-frame T2VA job go from 14922.98 MiB peak and 79.94 s to 14616.25 MiB and 82.27 s. That's 306.73 MiB off the peak, about 2%, and about 3% slower. The same doc is blunt that this isn't a universal memory or speed promise: gains depend on sequence length, quantization fusion, allocator state, and whatever else is on your GPU. If you're clearing 16GB with room to spare, this node costs you time for nothing.

Where it goes wrong

  • Conflicting wrappers. The node refuses to patch if another pack already owns the same H3 block/attention/MLP forward, if a DiT block replacement is installed via patches_replace["dit"], or if you connected the same T8 memory node twice. That failure is deliberate - it stops two algorithms from negotiating over one forward behind your back.
  • Don't stack KJ's equivalents. KJ's low-VRAM/ChunkFFN nodes and the H3 memory-efficient Sage node contend for the same modules. Pick one route.
  • Sage has a history. There's a well-documented case of global Sage corrupting output on a different architecture (patchy, blurred, matrix-looking lines or full black, reproduced on 3090/4090/5090). The safe pattern generally is to leave the global flag off and enable it per workflow through a node - which is what this node plays nicely with.
  • Red nodes everywhere means your ComfyUI core, frontend and Manager are behind. Update all three and do a full restart; updating the plugin alone typically doesn't fix it.

The MiniMax H3 Community License still applies to the weights, and it excludes the US, EU, UK and South Korea.

CategoryT8/MiniMax H3/Models/Experimental

Inputs (2)

NameTypeDefaultDescription
modelMODEL—
head_chunksINT41–56注意力头分组数。更大通常降低单次 attention 临时显存,但增加调用开销;1 仍启用提前释放,只是不分组。

Outputs (2)

NameTypeDescription
modelMODEL—
report_jsonSTRING—