Nodes/ComfyUI CV/CV Align MTB
ComfyUI Node

CV Align MTB

The first thing every hand-held exposure bracket needs

By bmad4ever·Created 4 months ago·Updated 14 days ago· 1
CV Align MTB
  • image
  • image
◄max_bits6►
◄exclude_range4►
◄cuttrue►

Stack three or five shots of the same scene at different shutter speeds and fuse them, and you get a photo with a sky that isn't blown and shadows that aren't mud. Do it hand-held and every frame is a few pixels off from its neighbours, so the fusion produces double edges on everything - the classic giveaway of a badly made HDR. CV Align MTB is the registration step: feed it the bracket as one batch, get the same frames back globally shifted so they line up.

Why you'd reach for it

It lives in the pack's hdr.py module, alongside the Mertens exposure fusion that's the actual reason to align: bracket → align → fuse. But translation-only registration is useful well beyond HDR:

  • Stabilise a shaky clip. A batch of near-identical video frames, each nudged back toward the common frame, kills hand-held jitter and nothing else.
  • Pre-align frames before a difference or median operation. Any "compare frames pixel by pixel" idea needs the frames registered first, and a global shift is often all it takes.
  • The pack's own docs are blunt that it's the cheap, honest version of registration: it corrects translation, period. No rotation, no scale, no perspective. If your camera rolled, this won't help and you want a homography/ECC route.

How it works

MTB - Median Threshold Bitmap - is Ward's trick for aligning exposures without ever comparing brightness values, which is the problem with naive alignment: the same pixel in a dark frame and a bright frame has wildly different values, so a difference metric is meaningless. Instead each frame is converted to grayscale and turned into a bitmap of "brighter than this frame's median" and "darker". The median is invariant to exposure, so the two bitmaps are directly comparable, and the best integer shift is found by searching a pyramid of XOR differences - pure integer logic, fast, no gradients, no brightness normalisation.

Three knobs, all about making that bitmap robust:

  • max_bits - the pyramid depth for the median bitmaps, default 6. Higher captures finer intensity variation and more noise. Leave it.
  • exclude_range - half-width of the pixel-value band around the median that gets excluded from the bitmap, default 4. Pixels near the median are exactly the ones that flip on sensor noise, so excluding them is noise rejection.
  • cut - drop the darkest and brightest pixels from the bitmap entirely, on by default. Those are the saturated highlights and crushed shadows, where the exposure relationship isn't a simple scale anyway.

Input image: a batch of 2+ frames - IMAGE, MASK, NPARRAY (list or 4-D array) or LATENT, because the socket is the pack's format-echoing match type. Output image: the aligned frames in the same format as the input. Stack your bracket with core Batch Images first; fewer than two frames and the node raises with a message telling you to do exactly that.

The bit you'd otherwise file a bug report about

OpenCV's AlignMTB.process() in the pinned 5.0.0.93 build zeroes the output slot of its internal reference frame - one frame can come back solid black from raw cv2. The pack detects that (a frame that came out all zeros) and restores the original, and it also converts single-channel and float frames to the 3-channel uint8 process() insists on, then converts them back. If you've ever cursed MTB alignment for a black frame, that's why the node has a guard in it.

Installing it

ComfyUI Manager → ComfyUI CV, or:

cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv

then restart ComfyUI. Requires opencv-contrib-python-headless~=5.0.0.93 (numpy and torch come with ComfyUI itself), Python 3.12+, and a recent ComfyUI with the V3 node API - on an older build the pack won't even import, and the nodes just don't appear. No model files involved: this is one of the ~317 hand-written nodes in a pack that's otherwise ~470 auto-generated cv2.* wrappers.

Traps

  • Two frames is the floor, and MTB likes more. With exactly two frames, "the median threshold bitmap" is a thin basis for a shift estimate; brackets of 3–7 are where it shines.
  • It shifts; it does not crop. A frame shifted by 20 pixels leaves a 20-pixel hole at one edge. In HDR fusion that's usually invisible after Mertens weights the overlap, but if you go on to compare frames exactly, you'll see the band.
  • Aligned does not mean identical. MTB optimises a bitmap match, not a photometric one. Rotate the tripod head half a degree and the shift it finds will be a compromise that helps nothing.
  • Don't align frames that are supposed to differ. If you're fusing a bracket, you want the exposures different and the geometry identical; if you're stacking frames for noise reduction, you want the opposite of that and near-identical exposure. Mixing the two gives you a smeared result and no obvious culprit.
Categoryimage/CV/hdr

Inputs (4)

NameTypeDefaultDescription
imageCOMFY_MATCHTYPE_V3Batch of 2+ images to align. Accepts IMAGE, MASK, NPARRAY (list or 4-D array), or LATENT. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size.
max_bitsINT61–32Maximum bit depth for the median bitmaps. 6 is typical; higher captures finer intensity variation but also more noise.
exclude_rangeINT40–128Half-width of the pixel-value range excluded from the median threshold bitmap. Helps reject noise.
cutBOOLEANtrueExclude the darkest and brightest pixels from the median threshold bitmap.

Outputs (1)

NameTypeDescription
imageCOMFY_MATCHTYPE_V3Aligned images in the same format as the input.