MiniMax H3 Fold AdaLN
MiniMax H3 Full-Width AdaLN, Folded Down to Rank 8
- output_path
- report
Two shapes of the same model
MiniMax H3 checkpoints show up in two layouts, and they are not interchangeable. ComfyUI's own H3 checkpoints ship the AdaLN modulation linears already folded - a rank-8 basis of the timestep-embedding curve plus a small adaln_t_table - and the time embedder is gone, because nothing needs it any more. The HuggingFace/diffusers-style full model instead carries full-width AdaLN linears of shape [*, 2688] and a live time_embedder.
That difference bites the moment you stop using stock weights. Fine-tunes, merges, dequantized originals - plenty come out full width, and the native ComfyUI path expects the folded form. This node folds them the same way ComfyUI's checkpoints are folded: every AdaLN linear onto the curve basis, time embedder swapped for adaln_t_table, saved.
Worth saying plainly: this is a file-layout conversion, not a licensing escape hatch. The weights are under the MiniMax H3 Community License, which excludes the US, EU, UK and South Korea from its applicable territory. Converting a local file doesn't change what you're licensed to run.
Why folding is even legal
AdaLN projection is a modulation signal. It takes the timestep embedding and produces per-layer shift/scale/gate vectors - it never sees arbitrary activations, only vectors that sit on the curve traced out by silu(time_embedder(t)) as t sweeps from 0 to 1. A full-width matrix applied only to on-curve inputs only ever explores a low-rank subspace of its output space. So sample that curve and find the subspace directly.
That's what the node does: load the time embedder's four tensors (proj_in weight and bias, proj_out weight and bias), sample silu(time_embedder(t)) across a grid of timesteps, SVD the result, keep the top curve_rank directions. Each AdaLN weight W is projected onto that basis (W @ basis), turning a [*, 2688] linear into a narrow one, while the table stores how to reconstruct the modulation vector for any timestep. The report gives the reconstruction error as a relative number, so you're not guessing how much was lost. Lossy, but bounded and measured.
Two behaviours that surprise people. Folded weights are cast to fp16 on the way out regardless of what came in - the projection runs in fp32 and lands in half precision. And the time embedder is only dropped when every AdaLN layer got folded; exclude some layers and it stays, because those kept-wide layers still need it. The node also patches the config metadata and stamps a modelutils_adaln_fold marker.
The inputs that matter
model_name and output_filename are the two you always set: input from ComfyUI's diffusion_models folder, output written back there (default minimax_h3_folded, no extension needed).
curve_rank (default 8, up to 64) and curve_grid (default 1025 points across t ∈ [0, 1]) match the ComfyUI H3 checkpoints - leave them alone unless you're trading fidelity for size. process_device picks where the SVD and projections run; cuda, with an automatic retry on CPU if you OOM. A 33B model's worth of folding on CPU is slow, but it finishes.
Then the filter trio, where the real control lives. discard_patterns omits tensors from the saved model entirely - dropping token_refiner blocks is the obvious use. exclude_patterns keeps matching AdaLN layers at full width, and include_mode flips that into a whitelist, so only matching layers fold (an empty include filter folds nothing). glob_patterns switches from Python regex substrings, the default, to shell globs. Patterns split on whitespace, so newline-separate them. discard_patterns beats everything else.
Outputs are output_path (the written file's name inside diffusion_models) and report: reconstruction error, layers folded, tensors preserved, tensors discarded, time-embedder tensors dropped, final size. It's an output node - nothing to connect.
Installing it
Registry name is Model Utility Toolkit; search that (or ComfyUI-ModelUtils) in ComfyUI Manager. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/silveroxides/ComfyUI-ModelUtils
cd ComfyUI-ModelUtils && pip install -r requirements.txt
That installs unifiedefficientloader>=0.5.2 (the streaming reader doing the work here), plus requests, Pillow, mutagen and av. The node uses comfy_api.latest, so an old ComfyUI will fail on import. The README doesn't document the MiniMax nodes yet - search the node menu.
Troubleshooting
Model is missing required time_embedder projection tensors. Not a bug, an answer: the file has noproj_in/proj_out, so it's already folded. Nothing to do here.- A shape error naming a key and its dimensions. An AdaLN
weightdidn't have the expected input width - you're either pointing at a non-H3 model or at an already-folded checkpoint. The error prints the actual shape, so you can tell which. - Nothing folded, everything preserved. Check
include_mode. With it on, only layers matchingexclude_patternsfold, and an empty pattern folds nothing at all. - Bad regex. Patterns are compiled and validated up front, so a typo raises immediately instead of silently matching nothing. Use
glob_patternsif you'd rather think in wildcards. - Low-bit inputs get rejected. The pack's guard refuses quantized tensors and quant sidecars rather than folding something it can't represent. Fold the full-precision file, quantize after.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | MiniMax H3 diffusion model with full-width AdaLN linears [*, 2688] to fold. | |
| output_filename | STRING | minimax_h3_folded | Output filename without extension, written under ComfyUI's diffusion_models directory. |
| discard_patterns | STRING | Newline-separated regex or glob patterns for tensors to omit entirely from the saved model (e.g. drop specific transformer blocks or refiner). | |
| glob_patterns | BOOLEAN | false | When True, filter patterns use shell globs (* matches any sequence, dots are literal). When False (default), patterns are Python regex matched as substrings. |
| curve_rank | INT | 81–64 | Basis rank for the silu(time_embedder) curve (default 8, matching ComfyUI H3 checkpoints). |
| curve_grid | INT | 102565–4097 | Number of interpolation grid points sampled for timestep t in [0, 1] (default 1025). |
| process_device | COMBO | cuda | Device used for SVD curve decomposition and matrix projection; CUDA OOM automatically retries on CPU. |
| exclude_patterns | STRING | Newline-separated regex or glob patterns. Matching AdaLN layers remain unfolded at full width [*, 2688]. | |
| include_mode | BOOLEAN | false | When True, use exclude_patterns as a whitelist: only matching AdaLN layers are folded; all others remain full width. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| output_path | * | — |
| report | STRING | — |