ComfyUI Node

Depth

Depth Anything V2 in your graph, with the licence you didn't read

By FXTD-Studios·Created 8 months ago·Updated about 18 hours ago· 246
Depth
  • image
  • depth_map
◄model_sizeLarge (335M - Best)►
◄normalizetrue►
◄invertfalse►
◄blur_edges0.0►
◄use_gputrue►

Depth Anything V2 has been this ecosystem's default monocular depth estimator since 2024 - nothing has knocked it off, not even Depth Anything 3. So this node isn't a novelty; it's a wrapper around the model most people were already reaching for through a ControlNet preprocessor. The difference here is you get the depth map as a first-class image you can wire anywhere, and the author added the video-flicker handling that the raw model doesn't have.

What it does

Runs Depth Anything V2 on an image (or a batch of frames) and returns a 3-channel grayscale depth map: light is near, dark is far, or inverted if you ask. That's it - one input, one output, depth_map (IMAGE).

Configured, it's the difference between plausible defocus and a fake blur. The node description frames the intended use as feeding a depth-of-field node for "realistic defocus blur", and that's the honest use case: a Gaussian blur with a depth map driving the radius looks dramatically more like a lens than a blur with a painted mask.

Beyond that: depth for control hints, depth for parallax fake-3D, depth as a matte for atmospheric haze, depth as a mask input to Radiance's own Energy Mask. It's a pass, and once you have a pass you'll find uses.

The inputs

model_size is the one real decision. Large (335M) is the default and it's the quality ceiling of the affordable tier - depth-estimation.md is blunt that Large has been the daily driver since 2024 for exactly the quality-per-second reason. Small (25M) is for fast previews and for real-time-ish video work; nobody is using it to deliver.

Then the optionals: normalize (on) rescales depth to 0–1, and for video the node standardises frames across the batch - which is the flicker fix, because per-frame normalization means frame 12 and frame 13 can map the same object to slightly different greys and the whole clip shimmers. invert flips near/far. blur_edges (0–5) smooths depth discontinuities, which matters when the depth map is driving a blur and you don't want a hard seam along a subject's edge. use_gpu (on) runs on CUDA and falls back to CPU when there isn't any.

One input behaviour from the tooltip worth knowing: frames with values above 1.05 are Reinhard tone-mapped into range first, and alpha is ignored. So you don't have to pre-tone-map HDR frames, but you do lose your alpha channel.

Install

Radiance handles the install and the download. Manager → search Radiance → install, or:

cd ComfyUI/custom_nodes
git clone https://github.com/fxtd-studios/radiance.git
cd radiance
python -m pip install -r requirements.txt

Windows portable: pip with python_embeded\python.exe. The models come from Hugging Face on first use (depth-anything/Depth-Anything-V2-Small-hf, -Base-hf, -Large-hf), and the pack pins the revision so an upstream change can't silently alter every depth map you've made. If you run with downloads disabled (RADIANCE_ALLOW_DOWNLOADS=0), only cached models load - which is the setup you want on an airgapped box, and the setup that produces "model not installed" errors if you forget.

Where people get burned

The licence. Depth Anything V2's repository is explicit: Small is Apache 2.0, but Base, Large and Giant are CC-BY-NC-4.0 - non-commercial. The default here is Large. If you're using this on client work, either accept that you're using a non-commercial weight, or switch to Small and lose quality. Worth a look before you ship, not after.

Video is not video-native. The normalization gives you temporal consistency of scale; it does not make the model track a sequence, so you can still get crawling edges on textured surfaces. The community read on this is honest and slightly deflating: one commenter in the pack's own release thread said they only reach for it for depth maps and that it "definitely won't create real 16/32 bit exr" - meaning depth as a real float pass with proper precision is a different problem from depth as a preview. For serious de-flickering you want a video-native depth model.

VRAM at resolution. Depth Anything V2 is a ViT; large resolutions are the cost driver, not the batch size. If it falls over at 4K, run it at half res and upscale the map - depth is a low-frequency signal and survives that badly-written move surprisingly well.

CategoryFXTD STUDIOS/Radiance/VFX

Inputs (6)

NameTypeDefaultDescription
imageIMAGEDisplay-encoded image or frame batch. Frames with values above 1.05 are Reinhard tone-mapped to 0..1 first; alpha is ignored.
model_sizeCOMBOLarge (335M - Best)Depth Anything V2 model size. Small = fast previews, Large = best quality.
normalizeoptBOOLEANtrueNormalize depth to 0-1 range. For video, frames are standardized for temporal consistency.
invertoptBOOLEANfalseInvert depth (white=far, black=near).
blur_edgesoptFLOAT0.00–5Gaussian blur to smooth depth discontinuities.
use_gpuoptBOOLEANtrueRun depth estimation on GPU. Requires a CUDA-capable device. Falls back to CPU if unavailable.

Outputs (1)

NameTypeDescription
depth_mapIMAGE—