CV Rasterize Mesh
The silhouette and depth you actually wanted from a 3D render
- pts3d
- tris
- K
- rvec
- tvec
- silhouette
- depth
- coverage
Most "render a mesh in ComfyUI" nodes give you a picture. This one gives you the two things you can actually compute with: a mask where the mesh covers the frame, and metric camera-space depth. It's a software z-buffer rasterizer, it needs no GPU and no model file, and - the reason it exists - it shares its z-buffer with this pack's CV Preview 3D (Calibrated Camera), so the matte you composite with and the picture you look at can't drift apart.
Why you'd reach for it
Think of the concrete jobs. You've tracked an object through a clip with CV Rapid Track (Sequence) and want a per-frame roto for inpainting or a ControlNet: that's a mask, per frame, from geometry rather than from a segmentation model that might decide the object is something else today. Or you want a real object in your plate to occlude a virtual one: that's the depth output, in the mesh's own units, which is exactly what the pack's 3D preview reads as scene_depth.
What it takes
pts3d (Nx3 vertices in OpenCV object space) and tris (Mx3 int triangle indices) come from CV Mesh From 3D Model, optionally through CV Mesh Split Long Edges or CV Transform Points 3D. The node renders geometry already in the graph - there's no model-file picker here, deliberately, so you can transform the mesh mid-graph.
K is the camera matrix, and it must be the same K the pose was solved with, or the silhouette lands in the wrong place. rvec / tvec take a single pose or the Bx3x1 stack that rapid tracking emits - that's the neat bit: one node turns a tracked clip into a per-frame matte and per-frame depth in a single execution. width / height set the output size; use the frame's own size so the mask lines up, and remember that if you halve the render you have to halve fx/fy/cx/cy in K too.
What comes out
silhouette is a MASK batch - one 0/1 mask per pose. It's hard-edged: there's no supersampling, so blur it if you need a soft matte, and expect stair-stepping on diagonal edges at small sizes.
depth is float32 camera-space Z in the mesh's units, HxW for a single pose and BxHxW for a stack. Pixels the mesh doesn't cover come back as inf, which is a deliberate choice rather than a nuisance: the pack's 3D preview reads inf as "no measurement here, never occlude". You can also sample it at projected points with CV Sample Array At Points to test whether an annotation anchor is actually facing the camera.
coverage is the mean fraction of the frame the mesh covers. Treat it as a health check, not a quality score: 0 means the model missed the frame entirely - wrong K, wrong units, or the pose sitting behind the camera - rather than that the render failed.
Install
Part of ComfyUI CV (bmad4ever/comfyui_cv), GPL-3.0, forked from opencv-comfyui. Manager: search the pack title. Manually:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
pip install "opencv-contrib-python-headless~=5.0.0.93"
# restart ComfyUI
Python ≥ 3.12 and a ComfyUI built on the V3 node API. This node is tagged NC in the pack's own docs - the rasterizer is render3d.py, first-party Python, not an OpenCV call. So the OpenCV version matters less here than it does for the wrapper nodes, but the rest of the pack still needs the contrib wheel.
Limitations, stated plainly
Pinhole only. Lens distortion is not applied. Undistort your frames first, or use the preview node when you need the lens modelled.
Triangles and silhouette only - no texture, no shading, no antialiasing. This emits geometry, not a picture. If you want the picture, that's CV Preview 3D (Calibrated Camera).
Triangle indices past the end of pts3d give you an empty render rather than reading out of bounds the way cv2.rapid.drawWireframe does. That's deliberate and it means a malformed mesh fails quietly - a blank mask is your bug report.
The best walkthrough is the pack's workflows/42_rapid_model_tracking.json, where the same pose stack drives a matte, a depth-based occlusion test, and annotations. If you only run one example from this pack, that's the one that shows how the 3D half fits together.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| pts3d | NPARRAY | Nx3 vertices in OPENCV object space, from 'CV Mesh From 3D Model' (optionally through 'CV Mesh Split Long Edges' or 'CV Transform Points 3D'). | |
| tris | NPARRAY | Mx3 int triangle indices into pts3d. Indices past the end of pts3d yield an empty render rather than reading out of bounds the way cv2.rapid.drawWireframe does. | |
| K | NPARRAY | 3x3 camera matrix of the frames being matched. Must be the SAME K the pose was solved with, or the silhouette lands in the wrong place. | |
| rvec | NPARRAY | Rotation, Rodrigues 3x1 - or a Bx3x1 STACK to render a whole tracked sequence in one execution. | |
| tvec | NPARRAY | Translation 3x1, or the matching Bx3x1 stack. Its units are the units the depth output is in. | |
| width | INT | 8008–8192 | Output width in pixels. Use the frame's own size so the mask lines up with the footage; K must match it (halving the render means halving fx/fy/cx/cy too). |
| height | INT | 6008–8192 | Output height in pixels, matching the frames. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| silhouette | MASK | One 0/1 MASK per pose (a MASK batch): 1 where the mesh covers the frame. Hard-edged - there is no supersampling here; blur it if you need a soft matte. This is the per-frame roto of a tracked object, ready for compositing, inpainting or a ControlNet. |
| depth | NPARRAY | METRIC camera-space Z in the mesh's own units, float32, and `inf` where the mesh does not cover the pixel - which is exactly what 'CV Preview 3D (Calibrated Camera)'.scene_depth reads as 'no measurement here, never occlude'. Shape is HxW for a single pose and BxHxW for a pose stack; take one frame out with 'CV Index Batch'. Sample it at projected points ('CV Sample Array At Points') to test whether an annotation anchor is facing the camera. |
| coverage | FLOAT | Mean fraction of the frame the mesh covers, over all poses. A health signal: 0 means the model missed the frame entirely (wrong K, wrong units, pose behind the camera), not that the render failed. |