CV 3D Object Transform
Place a model once, in the viewer and the tracker both
- object_transform
- matrix
Here's a specific, ugly problem in augmented-reality-style ComfyUI graphs. You have a 3-D model, and you want it in two places at once:
- Rendered over your photo or calibration target, by
CV Preview 3D (Calibrated Camera), which takes anobject_transformJSON string. - Tracked, meaning the actual geometry moved into the scene, by
CV Transform Points 3Dwith a 4×4 matrix - and the same placement has to driveCV Rapid Track, PPF, ICP,CV Rasterize Mesh.
Those two lanes live in different coordinate systems. The viewer is Three.js (Y up). OpenCV is Y down. So if you type the numbers in twice, you get a render and a tracker that disagree about where the object is, and the debugging is miserable because both look individually plausible.
This node is one set of numbers, two outputs, conversion handled.
The distinction the node is careful about
space - OpenCV (Y down, Z into the scene) or glTF / Three.js (Y up) - declares which convention your numbers are written in, and it should match CV Mesh From 3D Model's axis_convention. It is not a switch that moves the model. The two spaces are the same scene under a 180° roll about X, so typing y/z with the opposite sign in the other space gives an identical result. Flipping space alone therefore does move the model, but only because you changed the meaning of the numbers you typed - it's a promise, not an action. If you want the model to turn, use model_up and the rotations.
model_up is the action, and it's the dropdown that replaces the rotate_x_deg = -90 people type by hand in AR workflows: leave as the file has it, +Y is up (glTF / most GLB exports), +Z is up (Blender / CAD / OBJ). Note what "up" means here - it's the target plane's normal, OpenCV +Z, since the markers lie in its XY plane, not +Y. A stock +Y-up GLB arrives lying flat across the board, which is exactly why this widget exists. It's applied before your rotate_* values so those stay a free nudge on top.
The rest is a placement: position_x/y/z (same units as the mesh and the camera pose - metres if your calibration is), rotate_x_deg / rotate_y_deg / rotate_z_deg (each ±360, about the mesh origin - use recenter on CV Mesh From 3D Model if you want to spin about the object's own centre), and scale, which reconciles authoring units with the calibration. Rotation order is Rz·Ry·Rx, matching the pack's "CV Transform Points (Scale, Rotate, Translate)" blueprint.
Outputs
object_transform- the JSON string forCV Preview 3D (Calibrated Camera)'sobject_transforminput: position, quaternion [x, y, z, w], scale, in Three.js world space. Convert that widget to an input and wire this in; the viewer's gizmo then only previews, because this value wins on the next run.matrix- the same placement as a 4×4 float64 model matrix, in the selectedspace. Feed it toCV Transform Points 3Dto move the mesh, annotation anchors, or a cloud exactly as the render moved it.
CV Parse Object Transform is the return leg: read a placement dragged out with the viewer's gizmo back into a 4×4.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
Manager → ComfyUI CV. Restart and reload. Python ≥ 3.12, V3-API ComfyUI, opencv-contrib-python-headless~=5.0.0.93. No models needed for this node itself - though the workflow around it will want a GLB and a calibration.
Common issues
The render and the tracker disagreeing. Almost always space not matching the mesh's axis_convention. They're a pair; the tooltips say so and it's the one thing to check first.
An object at the right bearing and the wrong depth. A scale mismatch between authoring units and the calibration. A centimetre-authorised model against a metre calibration wants scale = 0.01, and the failure mode is directional rather than obvious.
A model standing on its side. model_up. Blender and CAD exports are usually +Z up; GLB is usually +Y up; a few are authored to stand already and want the default.
A gizmo drag that doesn't stick. By design. Once object_transform is wired from this node, the string wins every run - the gizmo becomes a preview. Use CV Parse Object Transform if you want the drag to become data.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| space | COMBO | OpenCV (Y down, Z into the scene) | Which convention the numbers below are written in, and the space 'matrix' comes back in - match it to 'CV Mesh From 3D Model'.axis_convention. The two are the same scene under a 180 degree roll about X (both right-handed), so this says how to READ the numbers below rather than what to do to the model: type y/z (and y/z rotations) with the opposite sign in the other space and the result is identical. Flipping this ALONE therefore does move the model - it is not a free switch, it is a promise about what your numbers mean. To turn the model on purpose, use 'model_up' and rotate_*. The 'object_transform' string is always Three.js whatever this says, because that is what the viewer reads. |
| model_up | COMBO | leave as the file has it | Which axis points UP inside the model FILE - a one-off turn that stands it on the calibration target, applied before your rotate_* values so those stay a free nudge on top. Note that UP here is the target's own normal (OpenCV +Z, since the markers lie in its XY plane), NOT +/-Y: a stock +Y-up GLB arrives lying flat across the board. Most GLB is +Y up (the manual equivalent is rotate_x_deg = -90); Blender/CAD/OBJ geometry is usually +Z up. Leave it alone for a model already authored to stand. Not the same job as 'space': that only renames a placement, this moves the model. |
| position_x | FLOAT | 0.00-1000000–1000000 | Translation along X, in the SAME units as the mesh and the camera pose (metres, if the calibration is). Applied last, after the rotation. |
| position_y | FLOAT | 0.00-1000000–1000000 | Translation along Y. This axis flips with 'space' - +Y is UP in Three.js, DOWN in OpenCV - so the same number moves the model opposite ways. |
| position_z | FLOAT | 0.00-1000000–1000000 | Translation along Z: OFF the target plane, since the board normal is the up direction here. +Z runs into the scene in OpenCV, toward the viewer in Three.js. |
| rotate_x_deg | FLOAT | 0-360–360 | Rotation about X, in degrees, about the mesh ORIGIN (use 'CV Mesh From 3D Model'.recenter to spin about the object's own centre). Applied FIRST of the three, and after 'model_up' - tilt, once the model already stands. |
| rotate_y_deg | FLOAT | 0-360–360 | Rotation about Y, in degrees. Applied second - the full order is Rz . Ry . Rx, matching the 'CV Transform Points (Scale, Rotate, Translate)' blueprint. |
| rotate_z_deg | FLOAT | 0-360–360 | Rotation about Z, in degrees. Applied LAST of the three - with 'model_up' set, this is the one that aims the standing model (its yaw about the board normal). |
| scale | FLOAT | 1.000.000001–1000000 | Uniform scale, applied before rotation and translation. Reconciles authoring units with the calibration (centimetres against a pose in metres needs 0.01) - a scale error reads as an object at the right bearing and the wrong depth. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| object_transform | STRING | JSON for the 'object_transform' input of 'CV Preview 3D (Calibrated Camera)': position, quaternion [x, y, z, w] and scale in Three.js world space. Convert that widget to an input and wire this into it; the viewer's gizmo then only previews, because this value wins on the next run. |
| matrix | NPARRAY | The same placement as a 4x4 float64 model matrix, in the selected 'space'. Feed it to 'CV Transform Points 3D' to move the mesh (or annotation anchors, or a point cloud) exactly the way the render moved the model. |