Latent → CV Array
Poking around inside the latent before it becomes pixels
- latents
- nparray
A ComfyUI LATENT isn't an image and doesn't pretend to be. It's a dict wrapping a tensor of shape [B,C,H,W], which means two things your usual node library can't touch: it's channel-first, and it carries a lot of channels - 4 for SD 1.5 and SDXL, 16 for Qwen-Image's VAE. Latent → CV Array unwraps it into a float ndarray [H,W,C] so the OpenCV side of the graph can actually look at it.
What comes out
latents takes the LATENT; batch_index picks which one of the batch to unwrap (clamped to the batch size, so a silly value quietly gives you the last frame rather than an exception). The single output nparray is float32 [H,W,C] with channels moved to the end.
The important word in the node's own description is untouched. This is not a decode. Nothing is quantised, nothing is normalised, no VAE is involved - you get the latent values as the sampler left them. That's what makes it the inverse of CV Array → Latent, and it's also why previewing the result directly looks like static. Latent values sit in a narrow range around zero, so if you want to see them, normalise or false-colour them first (CV Array → Mask does a min–max normalise, and CV Color Map renders a grayscale map as false colour).
Why you'd bother
Three real jobs. First, diagnostics: measure a latent instead of guessing - is the batch actually varying between seeds, or did the sampler converge to near-identical codes? Second, latent-space arithmetic that ComfyUI's core doesn't ship as nodes - an OpenCV call over latent channels is legitimate data work, just not photogenic. Third, handing latent data to something that only speaks NPARRAY, e.g. hashing, moments, or a low-level wrapper you're experimenting with.
Know the channel count before you write the next node. A [H,W,C] array with C=16 and a [H,W,C] with C=3 both look like images to cv2, and only one of them is. If you're on a 4-channel SDXL latent and you try to run a colour-space conversion, cv2 will oblige and produce numbers that mean nothing.
There's a batch-shaped sibling if you want the whole thing at once: Image Batch → CV Batch takes a LATENT and gives you [B,C,H,W] float32, untouched, no per-frame index needed.
Install
ComfyUI Manager, search the pack title comfyui_cv, repo bmad4ever/comfyui_cv. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart ComfyUI afterwards. The pack wants Python ≥ 3.12 and a recent ComfyUI on the V3 node API - it declares nodes with io.ComfyNode and io.Schema rather than NODE_CLASS_MAPPINGS, so an older ComfyUI just won't load it. Its one dependency:
pip install "opencv-contrib-python-headless~=5.0.0.93"
If you already have another OpenCV wheel installed, mind the shared-cv2 trap: all four distributions (opencv-python, opencv-python-headless, and both contrib builds) write to the same site-packages/cv2, and a non-contrib install over a contrib one silently removes the contrib submodules. Contrib nodes then disappear from the menu with no error in the log. The pack ships tools/repair_opencv_contrib.py with a --check and an --apply for that.
Where it goes wrong
The output looks like noise. It's a latent. It is noise, more or less. Normalise before you judge it.
Shape mismatch in the next node. [H,W,C] is not the same as the [B,H,W,C] you'd get from an IMAGE batch. If a node complains about dimensions, this is why, and there's no silent fix - reshape or index.
The latent is 16 channels and your filter assumes 3. Check nparray.shape conceptually before wiring a colour operation onto it.
The pack's README is worth ten minutes before you build on it: heavy LLM assistance in the code, test-driven overfitting to sample data in some workflows, unplanned updates, and an explicit "not recommended for production without independent review." Fair warning. For a one-line unwrap it hardly matters; for the DNN and matting nodes it matters a lot more.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| latents | LATENT | A LATENT dict with 'samples' tensor of shape [B,C,H,W]. The selected batch index is extracted and channels are moved to the last axis, yielding [H,W,C]. | |
| batch_index | INT | 00–4095 | Which latent of the batch to convert. Clamped to the batch size. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| nparray | NPARRAY | Float32 ndarray [H,W,C] (C = latent channels; 16 for qwen_image_vae, 4 for SD 1.5 / SDXL). |