ModelPack Diffusion Model
Fp8, Without Hunting Down the File
- MODEL
This is the workhorse of the pack: a UNet/DiT loader that takes an OCI artifact reference instead of a dropdown, and gives you one MODEL output. Same job as core's Load Diffusion Model - the difference is where the weights come from, and that you get to choose the precision on the way in.
Why you'd reach for it
Since Flux, the "checkpoint" is usually not one file. You've got a diffusion model, a text encoder, a VAE, and a workflow that assembles them - and the diffusion model is the 12-to-35 GB piece that people genuinely struggle to place correctly. Pointing this node at an artifact reference is one way to stop caring where that file lives: the artifact names it, the node pulls it, and the cache is under ComfyUI's own models directory.
It's also the pack's only precision control - nowhere else in it do you get a say in how weights are cast.
Inputs and output
reference- the OCI reference. Empty = local mode.file- path inside the artifact whenreferenceis set; otherwise a filename frommodels/diffusion_models. One of the two must be present.weight_dtype-default,fp8_e4m3fn,fp8_e4m3fn_fast,fp8_e5m2.
One output, MODEL - straight into a KSampler, or into a LoRA applier first.
What the four dtypes actually do
default loads the weights as they are on disk. That is the right answer more often than the menu implies: most "fp8" files on model hubs already contain fp8 weights, so loading them with default costs you exactly the same VRAM as casting. Pick it and move on unless you have a reason.
fp8_e4m3fn casts the weights to 8-bit with a 4-bit exponent and 3-bit mantissa - the standard fp8. This is the "99% identical to fp16 at half the VRAM" option the KB has been recommending for large models for two years, and it's the one to reach for when an fp16 file doesn't fit. Set your expectations correctly though: this happens at load time, so it saves VRAM, not download size. If you want a small artifact, the artifact has to ship fp8 weights in the first place.
fp8_e4m3fn_fast is the same dtype plus ComfyUI's fp8_optimizations flag turned on - the fast fp8 matmul path. On 40-series and newer hardware, which has fp8 acceleration, that's a real speed difference. On anything older, it isn't.
fp8_e5m2 gives the exponent five bits instead of four and the mantissa two: wider range, less precision. You'd reach for it when a model's outliers blow out e4m3, and most people never touch it.
The two mistakes people make with it
Casting fp8 on a 30-series card for speed. There's no fp8 acceleration there, so you get the VRAM saving and none of the throughput. That's not a reason to avoid it - it's a reason to know what you bought. If speed on a 3090/3060 is the actual goal, the current answer is native INT8 ConvRot, which ComfyUI has supported since v0.27.0 and which exists precisely because fp8 left older cards out.
Quantizing because you can, not because you must. The KB's blunt version of this: a 12 GB card holding a model that already fits at fp8 gains almost nothing from going smaller, and one A/B in the corpus found the real speedup came from removing --lowvram --reserve-vram launch flags rather than from quantizing at all. Load a pre-quantized file because the full-precision one doesn't fit. Don't cast when it already fits.
Installing it
Manager: search ComfyUI ModelPack (cerussite). Or:
cd ComfyUI/custom_nodes
git clone https://github.com/SiLeader/ComfyUI-ModelPack comfyui-modelpack
python -m pip install -r comfyui-modelpack/requirements.txt
Restart, and look under ModelPack/loaders. The only real dependency is modelpack-client (which brings oras, jsonschema, zstandard); nothing here rebuilds torch. Install into the Python that actually runs ComfyUI - on the Windows portable build that's python_embedded\python.exe. Python 3.10+ and ComfyUI 0.22.0+ are the floors.
Pull cache and failure modes
Weights land in ComfyUI/models/modelpack/diffusion_models/<hash>/, the hash being the first 24 characters of the SHA-256 of the reference. Use a digest reference (repo@sha256:…) and it's pulled once, then reused forever without contacting the registry; use a tag and it's re-pulled the first time it's used after each restart, with a fallback to the copy you already have if the registry is down. Change the reference string and you get a second copy - :v1 and :latest at identical bytes are not deduplicated. Private registries use your docker login credentials.
The node only checks that the pulled file has an extension ComfyUI can load; it has no idea whether the file is a diffusion model. Give it an artifact holding a VAE and the pull succeeds and the load fails, which looks confusing for about a minute. If you're not sure what's inside an artifact, leave file blank - the resulting error lists every weight path in the artifact, which is the closest thing to an inventory listing this pack offers.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| reference | STRING | OCI reference, e.g. registry.example.com/models/foo:v1. Leave empty to select a local file. | |
| file | STRING | Path within the artifact if it has multiple weights, or local ComfyUI model filename. | |
| weight_dtype | COMBO | 4 options: default, fp8_e4m3fn, fp8_e4m3fn_fast, fp8_e5m2 |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |