Nodes/ComfyUI-Marigold-v2/Marigold V2 Model Loader
ComfyUI Node

Marigold V2 Model Loader

The node that eats 41 GB so you don't have to think

By visualbruno·Created 29 days ago·Updated 21 days ago· 15
Marigold V2 Model Loader
    • marigold_model
    ◄modalitydepth►
    ◄depth_checkpointLog-stage2►
    ◄backboneQwen-Image-Edit-2509 bf16 (41 GB download)►
    ◄quantization4bit►
    ◄auto_downloadtrue►

    What this node actually is

    Marigold V2 reads a distance map, surface normals, or albedo off a single image. No prompt, no sampling loop, no CFG - one diffusion step through a frozen 20B Qwen-Image-Edit transformer with a rank-128 LoRA stacked on it. That's the whole trick, and it's why it can be a node in a graph instead of a script.

    Two ways to get that backbone in front of a sampler, and the pack ships a loader for each. This one is the self-contained path: it downloads the Qwen-Image-Edit-2509 backbone plus one ~1.8 GB checkpoint itself and runs them through diffusers, mirroring the reference scripts/infer.py. If you already have Qwen-Image-Edit weights, use Load MarigoldV2 LoRA instead - that path downloads nothing large, and it was about twice as fast with ~2 GB less VRAM in the author's own comparison of the two engines (Pearson r = 0.99959 on the same weights). So: not the performance pick. The "I don't own Qwen weights and I want the reference behavior" pick.

    How it works

    Qwen-Image-Edit minus the text encoder. The 7B encoder would be dead weight, so the checkpoints ship precomputed prompt embeddings instead - a few MB per modality, one set per modality value, and the one thing even the bring-your-own-weights path still fetches.

    The DiT is 20B and frozen. Each released checkpoint is a trainables.safetensors of roughly 1.8 GB holding the LoRA plus a fine-tuned VAE decoder. The backbone is shared across modalities, so once loaded, switching depth → normals → albedo swaps the adapter and the decoder in about a second.

    The inputs that matter

    modality - depth, normals, or albedo.

    depth_checkpoint - read only when modality is depth. Log-stage2 is the paper model; Log-stage1 is the pre-SinkLoss stage-1 model; Log-layered is see-through log depth for geometry behind glass; Uniform-base / Uniform-layered are linear depth, Marigold V1 style; Disparity-base / Disparity-layered are inverse depth. Log and linear grow with distance, disparity shrinks - the node knows which is which, so near_is_bright means the same thing on all seven. Normals and albedo have one checkpoint each and ignore this field.

    backbone - where the frozen DiT comes from. The bf16 entry is 41 GB and is the reference setup; the other six are the same 2509 model as GGUF, Q8_0 down to Q3_K_M (22 GB down to 10 GB). Any Qwen .gguf already sitting in models/diffusion_models, models/unet, or models/marigold-v2/gguf is offered by name too.

    quantization - 4bit (default, matches the released models), 8bit, or none. Only applies to the bf16 backbone; a GGUF is already quantized.

    auto_download - leave it on for the first run.

    Output is a single marigold_model, which goes into Marigold V2 Predict. No CLIP, no sampler, no merging.

    Backbone and VRAM, without the hand-waving

    The VAE and DiT config (~250 MB) download either way. Only the transformer weights differ, and that's where the decisions are.

    NF4 is more compact than Q4_K_M, so on a 16 GB card the 4-bit bf16 backbone actually leaves more room for activations than a Q4 GGUF does. The GGUF's advantage is download and load time: bf16 quantizes 20B parameters on first load and takes a few minutes, while a GGUF is up in under a minute. That's the opposite of the usual intuition. Q3_K_M at 10 GB is the smallest quant worth trying.

    Upstream's figures are ~17 GB at 1024×1024 and ~29 GB at 2048×2048 with the 4-bit DiT; quantization: none needs about 45 GB. If you're short, drop max_side on the Predict node rather than the checkpoint quality.

    The adapter was trained against 2509. A 2511 or 2512 GGUF does load and run - the DiT shape hasn't changed since - but that was checked on one image, not benchmarked, and the loader logs a warning when the filename doesn't say 2509.

    Install

    ComfyUI Manager, search ComfyUI-Marigold-v2. Or by hand:

    cd ComfyUI/custom_nodes
    git clone https://github.com/visualbruno/ComfyUI-Marigold-v2
    ../../python_embeded/python.exe -m pip install -r ComfyUI-Marigold-v2/requirements.txt
    

    Use your ComfyUI Python, not a system one. The README's own clone line still spells the author's older BrunoFargnoli handle - same repo.

    The dependency list is real: diffusers>=0.35 (for the Qwen-Image-Edit classes), peft, bitsandbytes, transformers, accelerate, safetensors, huggingface_hub, matplotlib. peft and bitsandbytes are the two missing from most installs. OpenCV is used if present and falls back to PIL/torch if not, so nothing force-installs a second cv2 over yours.

    Where the weights land, and the traps

    ComfyUI/models/marigold-v2/
    ├── Qwen-Image-Edit-2509/    vae/ (~0.25 GB), transformer/config.json, transformer/*.safetensors (~41 GB)
    ├── gguf/                    GGUF backbones fetched by these nodes
    └── Marigold-V2/             depth/Log-stage2/trainables.safetensors (~1.8 GB), normals/, albedo/,
                                 qwen_text_embeddings/
    

    Downloads resume if interrupted - re-run the workflow. To move the tree to another disk, point $MARIGOLD_V2_MODELS_DIR at it, or add marigold-v2 to extra_model_paths.yaml.

    Two things that look like crashes and aren't: the first bf16 load sitting silently for minutes (quantizing 20B parameters), and a run that appears to re-download 41 GB after an interruption (that's the resume working). If that wait sounds miserable, the LoRA node plus a Q4_K_M GGUF is the same model for a third of the traffic.

    One fair warning: there's no folklore to fall back on with this pack - as of the end of July 2026 the community corpus contains zero threads for "Marigold V2".

    CategoryMarigold V2

    Inputs (5)

    NameTypeDefaultDescription
    modalityCOMBOdepth3 options: depth, normals, albedo
    depth_checkpointCOMBOLog-stage2Depth parameterization; ignored for normals and albedo.
    backboneCOMBOQwen-Image-Edit-2509 bf16 (41 GB download)Where the frozen Qwen-Image-Edit DiT comes from. bf16 is the exact reference setup. A 2509 GGUF is the same model at a third of the download. A GGUF of another release (2511+) loads but was never what the adapter was trained against.
    quantizationCOMBO4bitApplies to the bf16 backbone only; GGUF weights are already quantized. 4bit matches the released models and needs ~17 GB of VRAM at 1024x1024, 'none' ~45 GB.
    auto_downloadBOOLEANtrueFetch missing weights from Hugging Face into ComfyUI/models/marigold-v2 (~41 GB backbone, ~1.8 GB per checkpoint).

    Outputs (1)

    NameTypeDescription
    marigold_modelMARIGOLDV2_MODEL—