Marigold V2 Model Loader
The node that eats 41 GB so you don't have to think
- marigold_model
What this node actually is
Marigold V2 reads a distance map, surface normals, or albedo off a single image. No prompt, no sampling loop, no CFG - one diffusion step through a frozen 20B Qwen-Image-Edit transformer with a rank-128 LoRA stacked on it. That's the whole trick, and it's why it can be a node in a graph instead of a script.
Two ways to get that backbone in front of a sampler, and the pack ships a loader for each. This one is the self-contained path: it downloads the Qwen-Image-Edit-2509 backbone plus one ~1.8 GB checkpoint itself and runs them through diffusers, mirroring the reference scripts/infer.py. If you already have Qwen-Image-Edit weights, use Load MarigoldV2 LoRA instead - that path downloads nothing large, and it was about twice as fast with ~2 GB less VRAM in the author's own comparison of the two engines (Pearson r = 0.99959 on the same weights). So: not the performance pick. The "I don't own Qwen weights and I want the reference behavior" pick.
How it works
Qwen-Image-Edit minus the text encoder. The 7B encoder would be dead weight, so the checkpoints ship precomputed prompt embeddings instead - a few MB per modality, one set per modality value, and the one thing even the bring-your-own-weights path still fetches.
The DiT is 20B and frozen. Each released checkpoint is a trainables.safetensors of roughly 1.8 GB holding the LoRA plus a fine-tuned VAE decoder. The backbone is shared across modalities, so once loaded, switching depth → normals → albedo swaps the adapter and the decoder in about a second.
The inputs that matter
modality - depth, normals, or albedo.
depth_checkpoint - read only when modality is depth. Log-stage2 is the paper model; Log-stage1 is the pre-SinkLoss stage-1 model; Log-layered is see-through log depth for geometry behind glass; Uniform-base / Uniform-layered are linear depth, Marigold V1 style; Disparity-base / Disparity-layered are inverse depth. Log and linear grow with distance, disparity shrinks - the node knows which is which, so near_is_bright means the same thing on all seven. Normals and albedo have one checkpoint each and ignore this field.
backbone - where the frozen DiT comes from. The bf16 entry is 41 GB and is the reference setup; the other six are the same 2509 model as GGUF, Q8_0 down to Q3_K_M (22 GB down to 10 GB). Any Qwen .gguf already sitting in models/diffusion_models, models/unet, or models/marigold-v2/gguf is offered by name too.
quantization - 4bit (default, matches the released models), 8bit, or none. Only applies to the bf16 backbone; a GGUF is already quantized.
auto_download - leave it on for the first run.
Output is a single marigold_model, which goes into Marigold V2 Predict. No CLIP, no sampler, no merging.
Backbone and VRAM, without the hand-waving
The VAE and DiT config (~250 MB) download either way. Only the transformer weights differ, and that's where the decisions are.
NF4 is more compact than Q4_K_M, so on a 16 GB card the 4-bit bf16 backbone actually leaves more room for activations than a Q4 GGUF does. The GGUF's advantage is download and load time: bf16 quantizes 20B parameters on first load and takes a few minutes, while a GGUF is up in under a minute. That's the opposite of the usual intuition. Q3_K_M at 10 GB is the smallest quant worth trying.
Upstream's figures are ~17 GB at 1024×1024 and ~29 GB at 2048×2048 with the 4-bit DiT; quantization: none needs about 45 GB. If you're short, drop max_side on the Predict node rather than the checkpoint quality.
The adapter was trained against 2509. A 2511 or 2512 GGUF does load and run - the DiT shape hasn't changed since - but that was checked on one image, not benchmarked, and the loader logs a warning when the filename doesn't say 2509.
Install
ComfyUI Manager, search ComfyUI-Marigold-v2. Or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/visualbruno/ComfyUI-Marigold-v2
../../python_embeded/python.exe -m pip install -r ComfyUI-Marigold-v2/requirements.txt
Use your ComfyUI Python, not a system one. The README's own clone line still spells the author's older BrunoFargnoli handle - same repo.
The dependency list is real: diffusers>=0.35 (for the Qwen-Image-Edit classes), peft, bitsandbytes, transformers, accelerate, safetensors, huggingface_hub, matplotlib. peft and bitsandbytes are the two missing from most installs. OpenCV is used if present and falls back to PIL/torch if not, so nothing force-installs a second cv2 over yours.
Where the weights land, and the traps
ComfyUI/models/marigold-v2/
├── Qwen-Image-Edit-2509/ vae/ (~0.25 GB), transformer/config.json, transformer/*.safetensors (~41 GB)
├── gguf/ GGUF backbones fetched by these nodes
└── Marigold-V2/ depth/Log-stage2/trainables.safetensors (~1.8 GB), normals/, albedo/,
qwen_text_embeddings/
Downloads resume if interrupted - re-run the workflow. To move the tree to another disk, point $MARIGOLD_V2_MODELS_DIR at it, or add marigold-v2 to extra_model_paths.yaml.
Two things that look like crashes and aren't: the first bf16 load sitting silently for minutes (quantizing 20B parameters), and a run that appears to re-download 41 GB after an interruption (that's the resume working). If that wait sounds miserable, the LoRA node plus a Q4_K_M GGUF is the same model for a third of the traffic.
One fair warning: there's no folklore to fall back on with this pack - as of the end of July 2026 the community corpus contains zero threads for "Marigold V2".
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| modality | COMBO | depth | 3 options: depth, normals, albedo |
| depth_checkpoint | COMBO | Log-stage2 | Depth parameterization; ignored for normals and albedo. |
| backbone | COMBO | Qwen-Image-Edit-2509 bf16 (41 GB download) | Where the frozen Qwen-Image-Edit DiT comes from. bf16 is the exact reference setup. A 2509 GGUF is the same model at a third of the download. A GGUF of another release (2511+) loads but was never what the adapter was trained against. |
| quantization | COMBO | 4bit | Applies to the bf16 backbone only; GGUF weights are already quantized. 4bit matches the released models and needs ~17 GB of VRAM at 1024x1024, 'none' ~45 GB. |
| auto_download | BOOLEAN | true | Fetch missing weights from Hugging Face into ComfyUI/models/marigold-v2 (~41 GB backbone, ~1.8 GB per checkpoint). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| marigold_model | MARIGOLDV2_MODEL | — |