ComfyUI Node

cv2.dct

JPEG's transform, exposed as two toggles

By bmad4ever·Created 4 months ago·Updated 15 days ago· 1
cv2.dct
  • src
  • nparray
◄flagsnone (0)►

The discrete cosine transform is the transform inside JPEG: it turns a block of pixels into frequency coefficients, which is why JPEG can throw away the high-frequency ones and still look like a photo. As a ComfyUI node it's a building block - you pair it with the inverse, with array math in this pack, or with cv2.xphoto.dctDenoising - rather than something you'd casually drop into an image pipeline.

What the node computes

cv2.dct takes a floating-point matrix and returns its DCT-II coefficients of the same size. Applied to an image it's a 2-D transform over the whole array; with the DCT_ROWS flag it transforms each row independently, which is the 1-D form you want when you're treating the data as a set of signals rather than a picture.

Two properties worth knowing before you wire anything. First, the transform pair is normalised - there's no SCALE flag in the DCT flag group (unlike the DFT, where people forget DFT_SCALE on the way back and get every value multiplied by the array size). Forward then inverse returns your original numbers. Second, OpenCV's DCT wants even-sized arrays, per its own documentation; feed it an odd width and you'll get an error or an implementation-defined surprise. Pad first - cv2.copyMakeBorder and cv2.getOptimalDFTSize are both in this pack - and crop afterwards if the size matters.

Inputs and outputs

There's a strict one: src accepts an NPARRAY only, and the tooltip says exactly why - "a data array (points / matrix), NOT an image". So you need Image → CV Array in front it (choose the float32 dtype option; the DCT refuses integer input) or another array-producing node. There's no way to wire a Load Image straight in, by design.

flags is a string widget rendered as a set of toggles, pipe-joined, e.g. none (0) | DCT_INVERSE. The DCT has only two, and the pack's own tooltip spells them out: DCT_ROWS transforms each row independently, and DCT_INVERSE turns this into the inverse transform - the same thing cv2.idct does. If you want to compose flags without remembering the spelling, this pack ships a CV DCT Flags builder node that emits exactly this string.

The output is a single NPARRAY of the same shape. To look at coefficients as a picture you need the usual dance: cv2.magnitude (or abs), a log to compress the enormous dynamic range, then CV Array → Image, which min-max normalises anything outside 0–255 back into viewable range.

Why you'd reach for it

Compression and reconstruction experiments - keep the low-frequency corner, zero the rest, inverse-transform, look at what survived. Feature vectors for image similarity, where DCT coefficients are a cheap fingerprint. Blocking-artifact and watermark work, since JPEG's DCT is the thing that causes 8×8 blocking in the first place. And if your actual goal is cleaning up compression noise, cv2.xphoto.dctDenoising in the same pack does the denoise-the-DCT-coefficients step for you rather than making you build it.

For a pure frequency-domain look at an image, cv2.dft is the more common door to walk through - the DFT is what everyone's Photoshop "FFT" plug-in uses, and the pack has a flag builder for it too. DCT is the one to pick when the domain is compression, not filtering.

Install

Manager → Install Custom Nodes → ComfyUI CV (publisher bmad4ever), or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"

Restart ComfyUI. Python ≥ 3.12 and a recent V3-API ComfyUI are required. The contrib headless wheel is the only dependency; dct itself is core OpenCV, so no model downloads. Installing a plain opencv-python wheel over the pack's contrib one overwrites the shared cv2 and silently empties the contrib submodules - python tools/repair_opencv_contrib.py --check and --apply handles that. The pack is GPL-3.0, forked from geroldmeisinger/opencv-comfyui, mostly LLM-authored, and the author states it is not production-ready and will not be supported promptly.

Common issues

  • "Unsupported depth" or an assertion about type. You fed integer data. Go through Image → CV Array with the float32 option, or cast with CV Cast Array.
  • cv2 raises on the array size. Odd dimensions. Pad to an even size with cv2.copyMakeBorder.
  • Round trip doesn't match the original. You mixed the 2-D transform with DCT_ROWS in one direction only. Both directions have to agree, and the same applies if you swapped cv2.idct for DCT_INVERSE mid-graph.
Categoryimage/CV/low-level/cv2 D

Inputs (2)

NameTypeDefaultDescription
srcNPARRAYinput floating-point array. A data array (points / matrix), NOT an image - only an NPARRAY link is accepted here.
flagsoptSTRINGnone (0)transformation flags as a combination of cv::DftFlags (DCT_*) cv2.dct flags: one of none (0) plus any of DCT_INVERSE, DCT_ROWS, pipe-joined (e.g. "none (0) | DCT_INVERSE"). In the UI this renders as a dropdown with one toggle per flag.

Outputs (1)

NameTypeDescription
nparrayNPARRAY—