Nodes/comfyui-minimax-h3-audio-T8/MiniMax H3 Chunk FeedForward / 分块前馈 (Advanced EXP/T8)
ComfyUI Node

MiniMax H3 Chunk FeedForward / 分块前馈 (Advanced EXP/T8)

The tiny patch you reach for when 16GB is barely enough

By T8mars·Created 2 months ago·Updated about 7 hours ago· 1,158
MiniMax H3 Chunk FeedForward / 分块前馈 (Advanced EXP/T8)
  • model
  • model
  • report_json
◄chunks2►
◄seq_threshold4096►

MiniMax H3 is a 33B joint video-and-audio model. On a 16GB card you are always one resolution bump, one extra reference frame, or one longer clip away from an out-of-memory error - and that is the whole reason this node exists. It will not make anything faster. It buys back a slice of peak VRAM by running H3's feed-forward network in token chunks instead of all at once, and it is honest about how small that slice is.

What it actually does

The node clones the MODEL you hand it and wraps the DiT MLP forward. When the packed token count is above your threshold, it splits the token axis into chunks pieces and runs the same SwiGLU on each - same equations, smaller activation peak per call, more GEMM launches. That last part is why it costs you time.

Nothing else changes. It doesn't touch attention, doesn't import KJNodes, and doesn't touch ComfyUI's process-global attention selection. Because it leaves attention alone, it can sit alongside any attention backend that doesn't replace the DiT block or MLP forward.

The inputs that matter

Three fields, one of which is the trap:

  • model - your H3 model after any ordinary weight LoRA.
  • chunks (default 2, 1–64) - how many pieces to split the token axis into. Higher usually means a lower activation peak, paid for in call overhead. At 1 the node is a genuine bypass: it returns the original MODEL object without cloning, checking, or patching anything. That's also the only bit-exact setting in the whole pack's memory feature set.
  • seq_threshold (default 4096, step 256) - chunking only kicks in when the packed token count is strictly greater than this. Read that again, because it's the field people miss: packed tokens scale with resolution × frame count, so a short clip can sit under 4096 and this node will do absolutely nothing. If you're memory-bound on short clips, lower it. If you want it out of the way on medium clips, raise it.

Outputs are model (wire it into your sampler, or into the Low VRAM Attention node, in either order) and report_json - a string with what actually got attached. Park it on a text node when something looks off.

Recommended wiring, from the pack's own notes:

H3 loader → optional weight LoRA → Low VRAM Attention → Chunk FeedForward → sampler

Install

ComfyUI Manager, search MiniMax H3 Audio T8. Or:

cd ComfyUI/custom_nodes
git clone https://github.com/T8mars/comfyui-minimax-h3-audio-T8.git minimax-h3-audio-T8

Restart ComfyUI fully afterwards. The pack's requirements.txt is deliberately empty - these nodes lean on the torch/numpy/Pillow that ComfyUI already ships, so there's nothing to pip install. You need a recent ComfyUI with native MiniMax H3 support (comfy.ldm.minimax), since the node validates the DiT block, attention and MLP signatures at load time.

What to expect, in numbers

The author's own fixed test - 3-second T2VA, 1024×512, 73 frames, 4 steps, same model/LoRA/prompt/seed - went from 14922.98 MiB peak / 79.94 s to 14616.25 MiB / 82.27 s with head_chunks=4 plus chunks=2. That is 306.73 MiB less peak (2.06%) and roughly 3% slower, and the two runs are different generations, not pixel-identical output. Useful if you're teetering on the OOM line; a waste of time if you're comfortable.

Where people get burned

The node fails closed on purpose, and the error messages are the feature. It refuses a second T8 memory node on the same branch, refuses when KJNodes or another pack already owns the same MLP forward, and refuses patches_replace["dit"] block replacements. If you see a rejection, don't hunt for a flag to force it - you have two things fighting over the same forward, and the pack is stopping you before the graph silently becomes a different algorithm. If you stack KJ's ChunkFFN or the H3 memory-efficient Sage node on top, this is what you'll hit.

Two more: the dual-model in-node long-video chain keeps its own MODEL identity whitelist and rejects these wrappers outright, so don't expect this node to work inside that workflow. And the usual pack-wide gotcha applies - if every node turns red or the workflow reports missing nodes, update ComfyUI itself, the frontend, and Manager, then do a full restart. Updating only the plugin is often not enough. One H3 job at a time on a 16GB card, always.

Licence note, since it is not optional: the MiniMax H3 Community License excludes the EU, UK, South Korea and the United States from its applicable territory - outputs included.

CategoryT8/MiniMax H3/Models/Experimental

Inputs (3)

NameTypeDefaultDescription
modelMODEL—
chunksINT21–64FFN 分块数。更大可能降低激活峰值,但会增加 GEMM 调用开销;1 返回原 MODEL。
seq_thresholdINT4096256–262144只有 packed token 数严格大于该值才分块;等于或低于阈值时保持单次原生 FFN。

Outputs (2)

NameTypeDescription
modelMODEL—
report_jsonSTRING—