Nodes/MiniMax H3 Activation Chunk - Star7/MiniMax H3 增强载入 - Star7
ComfyUI Node

MiniMax H3 增强载入 - Star7

The MiniMax H3 loader that keeps FP16 from cooking your video

By star7code·Created 2 months ago·Updated 5 days ago· 86
MiniMax H3 增强载入 - Star7
    • model
    ◄unet_name▾►

    MiniMax H3 is a 33B omni-modal model that generates video and its audio in one pass. If you only ever run it on a 4090 you may never need this node. If you're on an RTX 20-series, or you have --fp16-unet sitting in your launcher args out of habit, this is the loader that decides your precision policy before the sampler can do anything about it.

    What it is

    MiniMaxH3ChunkEnhancedLoaderStar7 (display name MiniMax H3 增强载入 - Star7) is a plain model loader - one dropdown, one MODEL output - with a precision policy baked in. It ships inside the MiniMax H3 Activation Chunk pack as the loader that pack expects upstream of its chunking node, and it uses its own class ID rather than reusing the standalone FP16 pack's loader, so both packs can sit installed without one shadowing the other.

    Reach for it when you're on Turing or older (RTX 20-series and below - anything under sm80), where there's no fast BF16 path and FP16 is your only performant option; or on a modern card when your launcher forces a global FP16 UNet. It does not chunk anything - the pack's chunk node does that - so don't install it expecting VRAM relief.

    How it picks your precision

    The gate is a device-capability check. No CUDA → default ComfyUI precision; ROCm → the FP16 path; sm61 → skipped, because Pascal's FP16 throughput is poor; sm80 and up → the model loads with bfloat16 dtype directly. That last one is the bit people miss: this loader overrides a global --fp16-unet setting on modern cards and hands you native BF16 with no repair wrappers attached.

    On the cards that get the FP16 path, the loader builds the model FP16 from construction - loading the state dict, detecting the H3 config, picking FP16 operations up front rather than casting after the fact. That matters for quantized H3 checkpoints (INT8, INT8+ConvRot), which route through MixedPrecisionOps: naively forcing FP16 dequantizes those layers and throws away the quantized kernels. The loader keeps force_cast_weights off so the INT8/ConvRot dispatch survives.

    Then it installs the overflow fix the name refers to. Every block's attn.out_proj and MLP fc2 forward gets wrapped so the projection runs on a pre-scaled input, the result is upcast to FP32, and the scale is multiplied back (64 and 256 respectively). Both are powers of two, so the rescale only moves the exponent - it's lossless, which is what "exact" is claiming. Net effect: the residual carrier that feeds the next block stays FP32 and can't overflow, while the heavy inner GEMMs stay in fast FP16. Without it, plain FP16 H3 on Turing is where you get the classic complaint set: NaNs, checkerboard patterning, garbage audio.

    The inputs and outputs

    There are exactly two:

    • unet_name - required, an enum populated from ComfyUI's diffusion_models folder. If the dropdown is empty, your H3 weights aren't in ComfyUI/models/diffusion_models (or an extra_model_paths entry pointing at it).
    • model (MODEL) - out to your LoRA loader, then the Activation Chunk node, then the guider.

    The file is validated: if you point it at a LoRA, a text encoder, or some other architecture, it fails with Selected file is not a native ComfyUI MiniMax H3 diffusion model rather than loading something wrong.

    Installing it

    ComfyUI Manager / Registry, searching the pack title:

    MiniMax H3 Activation Chunk - Star7
    

    Or by hand:

    cd ComfyUI/custom_nodes
    git clone https://github.com/star7code/minimax-h3-chunk-star7.git
    

    Restart ComfyUI afterwards. The pack pulls scipy, scenedetect>=0.7 and ultralytics>=8.3.162 - the last two belong to its video and face nodes, not to this loader. Once you're in the pack, the README's 20-series chain is:

    Enhanced Loader -> LoRA -> MiniMax H3 Activation Chunk - Star7 -> Guider
    

    Weights aren't included - H3 is a big download, and the Community License carves the US, EU, UK and South Korea out of the permitted territory, so check that first.

    Where it goes wrong

    The chunk node stops you before sampling on sm80+. If ComfyUI was launched with --fp16-unet, the chunk node refuses to run against unprotected FP16 and tells you to remove the flag or reload with the Star7 loader. Reloading here is the fix - you get BF16, which on a 30/40/50-series card is what you wanted anyway.

    A weight-patch warning in the log isn't noise. When the incoming model already carries weight patches - dynamic low-VRAM LoRA is the usual culprit - the loader notes that those patches can dequantize the INT8 layers it just worked to preserve, costing you some of the speedup.

    Don't stack the legacy patch on top. MiniMax H3 FP16 Exact Fix (Legacy) is a MODEL patch from the older design; the loader does that job at load time, which is strictly better, because the weights are created FP16 instead of cast later.

    It is not a VRAM fix. Weights, token count, sampler and resolution are untouched - if you're OOMing during sampling, that's the chunk node's job, and if you're OOMing at load time, that's a quantization decision.

    CategoryStar7/MiniMax H3

    Inputs (1)

    NameTypeDefaultDescription
    unet_nameCOMBO0 options:

    Outputs (1)

    NameTypeDescription
    modelMODEL—