Nodes/ComfyUI CV/CV Array → Latent
ComfyUI Node

CV Array → Latent

Pushing a cv2 array back into the sampler's world

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
CV Array → Latent
  • nparray
  • latent

What it's for

Feed a LATENT into any OpenCV node in this pack and it gets unwrapped: frame 0 becomes a plain float array so cv2 can touch it. That's great on the way in and a dead end on the way out, because every interesting thing downstream - VAE Decode, a sampler, an img2img pass - wants a LATENT, not an array.

CV Array → Latent is the way back. It wraps a float ndarray into the {'samples': [1, C, H, W]} structure ComfyUI's latent pipeline speaks, so a cv2-shaped result can re-enter the generative half of the graph.

The use case the author names is a good one: run CV Temporal Reduce over a LATENT batch to build a clean plate in latent space, wrap the plate with this node, and hand it to VAE Decode. You never decoded to pixels and re-encoded, so you never paid the round-trip loss. There's a whole example workflow for that shape of trick - exercise_background_subtraction_latent.json.

How it works

The array's last axis becomes the latent channel axis. A [H,W] float array becomes a single-channel latent; a [H,W,C] array becomes a C-channel one. Values are taken exactly as they are - no 0..1 scaling, no 0..255 scaling, and no quantisation, because latents are unbounded and clipping them would be a genuine information loss. The output is float32 and the batch dimension is fixed at 1.

The channel count is the part you have to get right yourself, and it's not guessable. A latent's channel count is a property of the specific VAE. The node's own note gives the two common reference points: 4 for the SD 1.5 / SDXL family, 16 for qwen_image_vae. Those numbers are not a rule about "modern models" - the VAE is a separately trained, sometimes separately licensed artifact now, and the whole point of the VAE-swap scene is that these spaces are not interchangeable. Check what your decoder expects rather than assuming.

The shape you need to think in

Latent dimensions are in cells, not pixels. For an 8× VAE, a [128, 128, 4] array decodes to a 1024×1024 image. If your array came out of an OpenCV node that was operating on an unwrapped latent, that's already the units you have and this node preserves them. If it came from a pixel-space operation, rescale before you get here, or you'll decode something 8× too small.

The workflow is: latent in → unwrapped to [H,W,C] float by the pack → your cv2 work, staying in float and staying untouched in value → back to a latent here → decode. Every step of that chain preserves both the value range and the exact shape, which is the entire reason to do latent-space OpenCV work at all.

Install

ComfyUI Manager, search ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

Restart, then:

pip install "opencv-contrib-python-headless~=5.0.0.93"

Python ≥ 3.12 and a recent ComfyUI (V3 node API) are hard requirements.

Where people get burned

The channel count. Decoding a 4-channel latent with a 16-channel VAE - or the reverse - doesn't produce a gentle warning, it produces an error or, worse, noise that looks like a model failure. Count the channels of the latent you started from and keep them.

One frame, not a batch. The output is [1, C, H, W]. There's no shape to put a sequence of latents into, so this is a single-frame door. If you need a video latent, that's a different structure entirely.

Don't normalise it. Latents routinely carry values well outside 0..1, and "helpfully" scaling them to a nice range before wrapping changes what the decoder sees. The node deliberately doesn't rescale; don't add a rescale in front of it.

Two-dimensional arrays make single-channel latents, which most decoders won't accept. If you have a grey mask and you want it to influence generation, the MASK socket is the right destination, not latent space.

Latent-space is not pixel-space for masking or colour. A latent is a compressed representation: 8× smaller in each spatial axis for a standard VAE. Spatial operations that are correct at cell resolution can look wrong at pixel resolution, and vice versa. The pack will happily let you mix them; you're the one who has to keep it straight.

Categoryimage/CV/low-level

Inputs (1)

NameTypeDefaultDescription
nparrayNPARRAYFloat ndarray [H,W] or [H,W,C] (C = latent channels; a 2-D array becomes single-channel). Values are taken as-is - no 0..1 or 0..255 scaling. Use 'Inspect CV Data' to check the shape/dtype.

Outputs (1)

NameTypeDescription
latentLATENT—