CV Optical Flow (TV-L1)
The accuracy reference for optical flow, priced in seconds per megapixel
- frame_a
- frame_b
- init_flow
- flow
- magnitude
- angle
If you want the best motion field classic (non-learned) computer vision can produce from two frames, this is it. TV-L1 minimizes an energy whose data term uses an L1 norm - which tolerates illumination changes and occlusions far better than Farneback's least squares - and whose prior is total variation, which keeps motion discontinuities sharp instead of smearing them across the object boundary. That combination is why it's the number everyone reports against.
The bill arrives as seconds per megapixel. Expect roughly an order of magnitude slower than DIS. So: TV-L1 for offline quality work and ground-truth-ish reference fields; DIS for anything you'd call interactive. That's the author's own framing and it's the right one.
The parameters, ranked
Frames go in as frame_a and frame_b (IMAGE or MASK direct, frame 0 of a batch, or NPARRAY), grayscale-converted internally, with frame_b resized to frame_a when they differ.
Four required parameters:
lambda_(default 0.15) - how much the data term weighs against smoothness. Lower means smoother: the TV prior wins and you get flatter, more regularized flow. Higher follows the pixels harder and picks up noise. It reads backwards from what most people expect.warps(default 5) - warping steps per pyramid level, the main speed knob. More accurate, linearly slower.scales(default 5) - pyramid levels; more captures bigger motion, each one costs time.epsilon(default 0.01) - the stopping threshold. Tighter is more accurate and much slower, so this is the knob that turns a 5-second node into a 60-second one.
Optional ones with sensible defaults: tau (0.25; the theory wants < 0.125 for guaranteed convergence, cv2 ships 0.25 because it works in practice), theta (0.3, the coupling tightness of the dual scheme), gamma (0, i.e. off - raise it to weight a gradient-constancy term when brightness changes between frames), scale_step (0.8), inner_iterations (30), outer_iterations (10), and median_filtering (5; set 1 to disable the outlier rejection).
And the useful one: init_flow, the previous frame pair's flow, to warm-start a video sequence. Copied before use, so an upstream cached array stays intact.
Outputs
flow (HxWx2 float32 (dx, dy)), magnitude (HxW float32, pixels), angle (HxW float32, radians - plugs straight into CV Flow To Color). Identical trio to every other flow node in the pack, so you can A/B TV-L1 against DIS by swapping one node.
Install
Manager → ComfyUI CV, or:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
cd comfyui_cv && pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart. Python ≥ 3.12, V3 node API. DualTVL1OpticalFlow is in cv2.optflow, so contrib is required - and because it's a class, no generated raw cv2_* wrapper can reach it, which is exactly why this curated node exists.
Common issues
- It takes forever - that's the algorithm, but check
epsilonandwarpsfirst; those two dominate. Also: on a long clip, run it on downscaled frames and upscale the flow field, because the cost is per-pixel. - Motion boundaries are soft - your
lambda_is too low. Raise it. - Everything is noisy and speckled -
lambda_too high. Lower it. - An occlusion or a flash breaks the field - set
gammaabove 0 to add brightness robustness rather than paying for a slower algorithm. - Instant failure mentioning
optflow- non-contrib OpenCV, or a plain wheel overwrote the contrib one insite-packages/cv2. The pack'stools/repair_opencv_contrib.py --checkreports it. DLL load failed while importing cv2on Windows portable - the multi-pack wheel collision other ComfyUI users report; it disables every cv2-based node at once. Not a graph problem.
Standard caveat for this pack: heavy LLM-assisted development, explicitly not production-ready, updates not promised. Here the risk is low - the node is a thin, faithful wrapper on OpenCV's own implementation.
Inputs (14)
| Name | Type | Default | Description |
|---|---|---|---|
| frame_a | NPARRAY,IMAGE | First frame (earlier in time); converted to grayscale. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| frame_b | NPARRAY,IMAGE | Second frame; resized to frame_a if the sizes differ. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| lambda_ | FLOAT | 0.150.001–2 | Weight of the data term against the smoothness term. LOWER = smoother flow (the prior wins); higher follows the pixels and picks up noise. |
| scales | INT | 51–10 | Pyramid levels. More levels capture larger motion; each costs time. |
| warps | INT | 51–20 | Warping steps per pyramid level - the outer linearization loop. More is more accurate and linearly slower; this is the main speed knob. |
| epsilon | FLOAT | 0.0100.0001–1 | Stopping threshold. Tighter (smaller) is more accurate and much slower. |
| init_flowopt | NPARRAY | Optional HxWx2 float32 flow to start from - pass the previous frame pair's flow when processing video. It is copied, never written to, so the upstream node's cached array stays safe. | |
| tauopt | FLOAT | 0.250.01–1 | Time step of the dual formulation. The theory needs tau < 0.125 for guaranteed convergence; cv2 ships 0.25 because it works in practice. |
| thetaopt | FLOAT | 0.300.01–2 | Tightness of the coupling between the two variables the dual scheme splits the problem into. Small values slow convergence. |
| gammaopt | FLOAT | 0.000–1 | Weight of the gradient-constancy term. Above 0 it adds robustness to brightness changes at some cost; cv2's default is 0 (off). |
| scale_stepopt | FLOAT | 0.800.1–0.95 | Size ratio between consecutive pyramid levels. |
| inner_iterationsopt | INT | 301–200 | Inner iterations used to solve the linearized problem at each warp. |
| outer_iterationsopt | INT | 101–100 | Outer iterations of the whole scheme. |
| median_filteringopt | INT | 51–9 | Aperture of the median filter applied between iterations to reject outliers. 1 disables it. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| flow | NPARRAY | HxWx2 float32 (dx, dy) displacement per pixel. |
| magnitude | NPARRAY | HxW float32 motion magnitude in pixels. |
| angle | NPARRAY | HxW float32 motion direction in RADIANS - feed 'CV Flow To Color' in its default radians mode. |