Marigold V2 Predict
One step, no CFG, no steps widget — that's not a bug
- marigold_model
- image
- vae
- image
- raw
What it does
This is the node that produces the map. It takes a loaded model, an image, and nothing else that resembles a sampling setting - because there's nothing to sample. Marigold V2 runs a single rectified-flow step at one fixed timestep, so there is no step count, no CFG, no scheduler, no denoise, and no prompt. If you went looking for the KSampler you're half-remembering, this is why it isn't here.
The context worth having: Marigold V1 was the slow one. It repurposed a Stable Diffusion model and denoised iteratively - roughly nine seconds a step, better with more steps and ensembling - which is why it never became anyone's daily depth preprocessor despite being sharp and generalising well to illustrations. V2 threw that away: one step, a 20B frozen transformer, a rank-128 adapter, three modalities out of one loaded backbone. For routine ControlNet work the knowledge base still says Depth Anything V2 Large; Marigold is the quality-and-oddball-input option that V2 finally made cheap.
The mechanism, briefly
Input goes through a Lanczos resize to a multiple of 16, the VAE encodes it, the DiT takes one step at t = 499/1000 with the checkpoint's precomputed prompt embedding standing in for the text encoder, the VAE decodes, and a per-modality adapter finishes the job: depth is averaged across the three channels into one grayscale plane, normals get unit-normalised into an RGB map, and albedo is converted out of linear RGB with a gamma-2.2 curve into sRGB. Depth comes out affine-invariant - meaningful up to an unknown scale and shift per image - which is why it's always normalised before you look at it.
Images are processed one at a time, deliberately. Upstream picks a different prompt context per batch position, which would make your result depend on where the image sat in the batch. A batch costs linearly and never runs out of memory because of its size.
The inputs that matter
marigold_model and image - the obvious two. image is a normal ComfyUI IMAGE, batch or single.
resolution_mode - native uses the input resolution rounded up to a multiple of 16, max_side downscales so the longer side is at most max_side, fixed gives you exactly width × height. So max_side is your VRAM dial, width/height only mean anything in fixed mode.
encoder_seed - seeds the VAE encoder's posterior sample, the only stochastic step in the whole pipeline. Default 2025, and its effect is small: two different seeds land at r = 0.99998, about 0.08% of the depth range. It advances per image in a batch, so the second image doesn't reuse the first's draw.
keep_input_size - on by default, resizes the prediction back to your original input resolution. The trip back down is area-averaged rather than bilinear, which matters if you fed it a big image or predicted above input resolution.
near_is_bright - depth only. Bright means close to camera. The node knows your checkpoint's convention (log and linear grow with distance, disparity shrinks) so this flag means the same thing whichever you loaded.
vae_tiling - trades a little quality at the seams for VRAM at high resolutions. One catch: it only exists on the diffusers engine. On the bring-your-own-weights path, turning it on logs a warning and does nothing - lower max_side instead. Same for keep_model_loaded, which only matters when the backbone came from Marigold V2 Model Loader.
vae (optional) - required when the model came from Load MarigoldV2 LoRA. Wire the LoRA node's vae output here, not a bare Load VAE.
Outputs
image is the finished, viewable thing: normalised grayscale for depth, an RGB normal map for normals, sRGB albedo for albedo. Preview it, save it, or feed a depth ControlNet - the depth variant is already normalised and near-is-bright, which is the layout a depth ControlNet expects.
raw is the float prediction in the model's own units - [H, W] for depth, [3, H, W] for normals and albedo. It's what Marigold V2 Colorize Depth and Marigold V2 Save Raw (.npy) consume. The image output is a rendering; the raw output is the data.
Install
ComfyUI Manager, search ComfyUI-Marigold-v2, or:
cd ComfyUI/custom_nodes
git clone https://github.com/visualbruno/ComfyUI-Marigold-v2
../../python_embeded/python.exe -m pip install -r ComfyUI-Marigold-v2/requirements.txt
Then drop an example_workflows/*.json onto the canvas - marigold_v2_depth.json for the self-contained path, marigold_v2_comfyui_native.json for the one that reuses your Qwen weights.## Where it goes wrong
Out-of-memory at 1024×1024 means you're near 17 GB of VRAM; upstream's figure at 2048×2048 is ~29 GB. In order of least painful: lower max_side, switch to a smaller GGUF backbone, or enable vae_tiling if you're on the diffusers path.
If depth looks inverted - sky blazing white and the foreground black - flip near_is_bright. That's the one setting people get wrong, because it's a display convention rather than a model property, and the raw values underneath don't change at all.
And if you pass a normals or albedo prediction into the colorize node you'll get a hard error rather than a weird-looking map. That's deliberate; the error message tells you which modality it actually got.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| marigold_model | MARIGOLDV2_MODEL | — | |
| image | IMAGE | — | |
| resolution_mode | COMBO | max_side | native: the input resolution, rounded up to a multiple of 16. max_side: downscale so the longer side is at most max_side. fixed: exactly width x height. |
| max_side | INT | 102416–8192 | — |
| width | INT | 102416–8192 | — |
| height | INT | 102416–8192 | — |
| encoder_seed | INT | 20250–18446744073709550000 | Seeds the VAE encoder sample, the only stochastic step. Its effect on the result is small. |
| keep_input_size | BOOLEAN | true | Resize the prediction back to the input resolution. |
| near_is_bright | BOOLEAN | true | Depth only: bright pixels are close to the camera. |
| vae_tiling | BOOLEAN | false | Trades quality for VRAM at high resolutions. |
| keep_model_loaded | BOOLEAN | true | Diffusers backbone only; the ComfyUI engine leaves loading and unloading to ComfyUI. |
| vaeopt | VAE | Required when the model came from Load MarigoldV2 LoRA. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| raw | MARIGOLDV2_RAW | — |