Image Batch → CV Batch
Run OpenCV on a whole clip instead of per-frame
- input
- batch
Most OpenCV-in-ComfyUI workflows are a for loop wearing a graph costume: convert frame, filter frame, convert back, repeat 120 times. Image Batch → CV Batch is the node that lets you skip the loop. It turns an entire ComfyUI batch into one batched ndarray - [B,H,W,C] - so a single cv2 call (or a single curated node that's written batch-aware) sees every frame at once.
Inputs and outputs, precisely
input is a multi-type socket: IMAGE, MASK, LATENT or NPARRAY, and it always takes the entire batch. There's no batch_index here - that's the whole point, and it's the difference between this node and Image → CV Array, which grabs one frame.
What you get in batch depends on what you fed it:
- IMAGE →
[B,H,W,C]in BGR, float32 by default. - MASK →
[B,H,W]float32 in 0–1. - LATENT →
[B,C,H,W]float32, values untouched (the scaling is ignored for latents, they're always float32). - NPARRAY → passes straight through.
dtype only affects IMAGE/MASK, and it's the one setting worth thinking about. The default, float32 (0-1, no quantisation), keeps the precision you already have: no [0,1] → [0,255] round trip, and - the part that bites - no overflow clamping when you add, multiply or mix images. The uint8 (0-255, standard cv2) option is for when a low-level cv2 wrapper really does want the classic 8-bit path.
The output is a raw NPARRAY: no type preservation, no memory of where it came from. To get back, wire it into CV Batch → Image Batch (or Latent → CV Array if you're heading back to latent space).
Where this earns its place
Anything frame-symmetric. Temporal reduce to a background plate, a per-frame threshold, a batch of hashes, a stack of HOG descriptors, quality-scoring every candidate in a batch against one reference. Doing it batched is both faster and dramatically less graph spaghetti than a frame-by-frame fan-out.
Two things to keep in mind. First, "the pack's nodes accept the batched shape" is a convention the curated nodes honour (the tooltips on several of them say arithmetic ops process the full IMAGE batch when both inputs share a batch size) - a thin auto-generated wrapper that wants a single 2-D image may not, and you'll get a shape error rather than a helpful message. Second, shape is data here: [B,H,W,C] from an IMAGE and [B,C,H,W] from a LATENT look nothing alike to a cv2 function, so know which one is on the wire before you write the next node.
Install
ComfyUI Manager, search comfyui_cv. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Restart. Requirements: Python ≥ 3.12, a recent ComfyUI on the V3 node API (this pack is written entirely in io.ComfyNode/io.Schema), and one Python dependency:
pip install "opencv-contrib-python-headless~=5.0.0.93"
The pin is the version the pack is curated against, and the contrib part is not optional - a non-contrib OpenCV wheel installed over a contrib one shares the same site-packages/cv2 and silently guts the contrib submodules, taking contrib nodes with it. tools/repair_opencv_contrib.py --check and --apply are in the pack for exactly that mess.
Getting burned, the usual ways
You expected [B,H,W,C] and got [B,C,H,W]. You fed it a LATENT. Channel-first is deliberate and identical to what the latent actually is; flip the axes mentally, or convert with Latent → CV Array if you want one frame as [H,W,C].
Downstream node rejects the batch. Some cv2 wrappers are single-image calls. If a batched input errors, take the batch apart with CV Index Batch and drive it one frame at a time - slower, but it always works.
Values came out wrong after arithmetic. Check dtype. On float32 you asked for no clamping; if something downstream assumed 0–255, that's the mismatch, not the maths.
One caveat from the README, since it applies to the whole pack: this was built with heavy LLM assistance, several example pipelines are overfitted to their sample data, and it's explicitly not production-grade without you reviewing the code. The conversion nodes are the least risky part of that - they do one boring thing - but don't assume the same for the 470 wrappers behind them.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| input | IMAGE,MASK,LATENT,NPARRAY | IMAGE, MASK, LATENT or NPARRAY batch to convert. The entire batch is extracted (no batch_index). | |
| dtype | COMBO | float32 (0-1, no quantisation) | For IMAGE/MASK: float32 keeps the original precision (no [0,1]→[0,255] quantisation and no overflow clamping on add/multiply); uint8 matches the standard cv2 path. For LATENT this is ignored (always float32). For NPARRAY the value passes through as-is. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| batch | NPARRAY | Batched ndarray: [B,H,W,C] from IMAGE, [B,H,W] from MASK, [B,C,H,W] from LATENT, or the NPARRAY as-is. |