cv2.dct
JPEG's transform, exposed as two toggles
- src
- nparray
The discrete cosine transform is the transform inside JPEG: it turns a block of pixels into frequency coefficients, which is why JPEG can throw away the high-frequency ones and still look like a photo. As a ComfyUI node it's a building block - you pair it with the inverse, with array math in this pack, or with cv2.xphoto.dctDenoising - rather than something you'd casually drop into an image pipeline.
What the node computes
cv2.dct takes a floating-point matrix and returns its DCT-II coefficients of the same size. Applied to an image it's a 2-D transform over the whole array; with the DCT_ROWS flag it transforms each row independently, which is the 1-D form you want when you're treating the data as a set of signals rather than a picture.
Two properties worth knowing before you wire anything. First, the transform pair is normalised - there's no SCALE flag in the DCT flag group (unlike the DFT, where people forget DFT_SCALE on the way back and get every value multiplied by the array size). Forward then inverse returns your original numbers. Second, OpenCV's DCT wants even-sized arrays, per its own documentation; feed it an odd width and you'll get an error or an implementation-defined surprise. Pad first - cv2.copyMakeBorder and cv2.getOptimalDFTSize are both in this pack - and crop afterwards if the size matters.
Inputs and outputs
There's a strict one: src accepts an NPARRAY only, and the tooltip says exactly why - "a data array (points / matrix), NOT an image". So you need Image → CV Array in front it (choose the float32 dtype option; the DCT refuses integer input) or another array-producing node. There's no way to wire a Load Image straight in, by design.
flags is a string widget rendered as a set of toggles, pipe-joined, e.g. none (0) | DCT_INVERSE. The DCT has only two, and the pack's own tooltip spells them out: DCT_ROWS transforms each row independently, and DCT_INVERSE turns this into the inverse transform - the same thing cv2.idct does. If you want to compose flags without remembering the spelling, this pack ships a CV DCT Flags builder node that emits exactly this string.
The output is a single NPARRAY of the same shape. To look at coefficients as a picture you need the usual dance: cv2.magnitude (or abs), a log to compress the enormous dynamic range, then CV Array → Image, which min-max normalises anything outside 0–255 back into viewable range.
Why you'd reach for it
Compression and reconstruction experiments - keep the low-frequency corner, zero the rest, inverse-transform, look at what survived. Feature vectors for image similarity, where DCT coefficients are a cheap fingerprint. Blocking-artifact and watermark work, since JPEG's DCT is the thing that causes 8×8 blocking in the first place. And if your actual goal is cleaning up compression noise, cv2.xphoto.dctDenoising in the same pack does the denoise-the-DCT-coefficients step for you rather than making you build it.
For a pure frequency-domain look at an image, cv2.dft is the more common door to walk through - the DFT is what everyone's Photoshop "FFT" plug-in uses, and the pack has a flag builder for it too. DCT is the one to pick when the domain is compression, not filtering.
Install
Manager → Install Custom Nodes → ComfyUI CV (publisher bmad4ever), or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart ComfyUI. Python ≥ 3.12 and a recent V3-API ComfyUI are required. The contrib headless wheel is the only dependency; dct itself is core OpenCV, so no model downloads. Installing a plain opencv-python wheel over the pack's contrib one overwrites the shared cv2 and silently empties the contrib submodules - python tools/repair_opencv_contrib.py --check and --apply handles that. The pack is GPL-3.0, forked from geroldmeisinger/opencv-comfyui, mostly LLM-authored, and the author states it is not production-ready and will not be supported promptly.
Common issues
- "Unsupported depth" or an assertion about type. You fed integer data. Go through
Image → CV Arraywith the float32 option, or cast withCV Cast Array. - cv2 raises on the array size. Odd dimensions. Pad to an even size with
cv2.copyMakeBorder. - Round trip doesn't match the original. You mixed the 2-D transform with
DCT_ROWSin one direction only. Both directions have to agree, and the same applies if you swappedcv2.idctforDCT_INVERSEmid-graph.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| src | NPARRAY | input floating-point array. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here. | |
| flagsopt | STRING | none (0) | transformation flags as a combination of cv::DftFlags (DCT_*) cv2.dct flags: one of none (0) plus any of DCT_INVERSE, DCT_ROWS, pipe-joined (e.g. "none (0) | DCT_INVERSE"). In the UI this renders as a dropdown with one toggle per flag. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| nparray | NPARRAY | — |