CV Stereo Calibrate (Chessboard)
Get R and T Right the First Time
- images_left
- images_right
- K_left
- dist_left
- K_right
- dist_right
- K_left
- dist_left
- K_right
- dist_right
- R
- T
- rms_error
- found
You have two cameras bolted to a bar. Everything you want downstream - real depth, real distances in millimeters, a point cloud you can trust - depends on one thing: where the right camera sits relative to the left. That's a rotation and a translation, and this node is where cv2.stereoCalibrate gets to estimate them.
If you've only ever done monocular depth (MiDaS, Depth Anything), here's the difference that matters: those models infer depth from a single image and give you relative ordering - nearer/farther, no units. A calibrated stereo rig measures it. Set square_size in real millimeters and your disparity map becomes metres.
How it works
You feed two IMAGE batches, left and right, of the same board poses in the same order. For each pair the node runs chessboard detection independently on both frames and refines the corners with cornerSubPix; a pair only counts if both views found the board. Object points come from a synthetic grid scaled by square_size. Then all the accumulated 3-D/2-D correspondences go into one cv2.stereoCalibrate call.
The wrinkle worth understanding is the intrinsics. OpenCV can solve for both cameras' intrinsics and the relative pose in one shot, and it can also hold the intrinsics fixed - and the node's behaviour depends entirely on what you wire in:
- Both
K_leftandK_rightconnected →CALIB_FIX_INTRINSIC. Only R and T are estimated. This is the recommended two-stage workflow: calibrate each camera on its own first, then pin those results here. - One or neither connected → the intrinsics are estimated or refined here, and any K you did pass is just an initial guess. This is the trap people walk into: connecting a single K does not hold it.
The individual pre-pass is CV Calibrate Camera (Chessboard), which is the one you usually want.
What you actually set
images_left and images_right must be the same length and the same order - pair 3 on the left has to be the same physical board pose as pair 3 on the right. Roughly 10–20 usable pairs gives a good fit; the node needs at least 3.
pattern_cols and pattern_rows are inner corners, not squares - a board with 10 × 7 squares is 9 × 6 here, hence the defaults. Off-by-one on this is the single most common reason the detector finds nothing at all.
square_size is the world scale. Leave it at 1.0 and everything comes out in arbitrary units, which is fine for rectification and useless for measuring.
On output: K_left, dist_left, K_right, dist_right, plus R and T (left-to-right pose), rms_error, and found. Wire R, T, K and dist into cv2_stereoRectify, remap both images, then feed the rectified pair to CV Stereo Disparity (SGBM). rms_error is the mean reprojection error in pixels - a fraction of a pixel is a good fit, and if you're staring at several pixels something upstream is wrong (board too flat in the frame, motion blur, or you're trying to calibrate pairs that aren't actually paired).
Install
ComfyUI Manager → search ComfyUI CV (publisher bmad4ever), or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
Restart after. The pack needs Python ≥ 3.12 and a recent ComfyUI on the V3 node API. The contrib wheel is not optional: install a plain opencv-python on top and it silently shares site-packages/cv2, wiping the contrib submodules and making a pile of nodes disappear. There's a tools/repair_opencv_contrib.py --check / --apply in the repo for exactly that mess.
The example photos come with the repo - run workflows/01_install_example_inputs.json once and reload the page, then look at 49_chessboard_playground.json and exercise_stereo_multiview.json.
Where it bites
The node is deliberately failure-tolerant: fewer than 3 usable views (or a pair set that never detected) returns found=false, an identity R and a zero T - no exception, no red node. Branch on found, not on rms_error, because the fallback reports an rms of 0.0.
Also be honest about the surrounding pack: bmad4ever's own README says the workflows are showcase-grade, with stereo settings tuned against StereoGeo-CARLA, and flags the whole codebase as heavily LLM-generated with no support promised. The node is a straight cv2.stereoCalibrate wrapper and behaves like one; the heuristics around it are what you should treat as a starting point.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| images_left | IMAGE | Batch of LEFT chessboard views (>= 3 usable pairs; ~10-20 gives a good fit). Must be in the same order as images_right. | |
| images_right | IMAGE | Batch of RIGHT chessboard views, same order/length as images_left. | |
| pattern_cols | INT | 92–40 | Inner corners per row (squares per row minus 1). |
| pattern_rows | INT | 62–40 | Inner corners per column (squares per column minus 1). |
| square_sizeopt | FLOAT | 1.000.0001–1000000 | Physical size of one square (e.g. mm); sets the world scale. Leave 1.0 for relative calibration. |
| K_leftopt | NPARRAY | 3x3 intrinsics for the left camera from 'Calibrate Camera'. Connect BOTH K_left and K_right to hold the intrinsics FIXED and solve only for R/T; on its own it is just an initial guess and still gets refined. Leave unconnected for a fresh estimate. | |
| dist_leftopt | NPARRAY | Left distortion coefficients from 'Calibrate Camera'. Also held fixed when both K matrices are connected. | |
| K_rightopt | NPARRAY | 3x3 intrinsics for the right camera. See K_left: both connected = intrinsics pinned, one alone = initial guess only. | |
| dist_rightopt | NPARRAY | Right distortion coefficients. |
Outputs (8)
| Name | Type | Description |
|---|---|---|
| K_left | NPARRAY | Refined 3x3 intrinsic matrix for the left camera. |
| dist_left | NPARRAY | Refined distortion coefficients for the left camera. |
| K_right | NPARRAY | Refined 3x3 intrinsic matrix for the right camera. |
| dist_right | NPARRAY | Refined distortion coefficients for the right camera. |
| R | NPARRAY | 3x3 rotation matrix from left to right camera. |
| T | NPARRAY | 3x1 translation vector from left to right camera. |
| rms_error | FLOAT | Mean reprojection error in pixels; lower is better. |
| found | BOOLEAN | — |