On-Set Studio
Build the shot first, then let the model paint it
- out_1
- out_2
- out_3
- out_4
- character_json
- environment_json
- scene_json
- slot_labels
- openpose_json
What this thing actually is
Every control workflow you've built works backwards. You generate or find an image, a preprocessor extracts depth, canny, an OpenPose skeleton, and a second model redraws it. Fine pipeline - but your composition was decided by whatever came out of the first sampler.
On-Set Studio flips it. It's a browser-based 3D editor - a virtual stage - running as a ComfyUI custom node. Import a Mixamo-rigged character, block the scene, rig lights, choose a lens, lay a camera track, press Send to ComfyUI. What lands in your graph is a set of passes rendered from that exact camera: scene depth, character-only depth, normals, canny, a beauty render, flat character ID mattes, an OpenPose control image, plus structured JSON.
So the control map isn't estimated from a picture. It's authored. Every pass renders from the same camera and framing, so depth and pose register by construction - no drift, which is the usual reason people give up stacking depth + pose + canny.
Built on the bones of 3D Openpose Editor (MIT, credited), version 1.0.2, one author, sharp edges. Nothing else does this job inside ComfyUI.
Where it fits: it's the A side. You stage, it hands you ground truth, and the rest of the graph is your normal stack. Per-frame control video models are the natural consumer, since a sequence send gives you a matched control clip rather than one frame.
How the handoff works
No APIs, no keys, no cloud round-trip. The editor is served locally by the pack at /on-set-studio/app/, and it writes a payload onto disk inside the pack's own folder:
- A still send writes
payloads/<session>.json. - A sequence send writes
payloads/<session>/manifest.jsonpluspayloads/<session>/frames/00000_<pass>.pngand friends.
The node takes the newest - whatever you pressed last wins, with no mode widget to keep in sync. A still arrives as a normal [1,H,W,3] image; a sequence arrives on the same socket as an [N,H,W,3] batch, which is what video nodes want. Flipping a slot from Frame to Video doesn't break a saved graph.
The four slots are routed in the editor's Settings tab, not on the node. Each picks a map from a dropdown (Depth: scene, Depth: character only, Normal: character, Canny: edges, Pose: OpenPose skeleton, Character mattes: flat ID, None: output disabled, and more). Sockets stay put because ComfyUI links by index - only the cargo changes.
The inputs and outputs that matter
One input: session, a string defaulting to default - the label on the payload folder. Change it before opening the editor if more than one tab is sending to the same ComfyUI.
Then four IMAGE outputs, out_1 to out_4. Those are what you wire. A slot set to None: output disabled emits no frame at all rather than a black one, which is what you want: a four-slot sequence send costs four renders per frame.
The rest are strings: character_json, environment_json, scene_json (per-character, set and scene-level data, including the BBOX data the editor generates), slot_labels - which map is on which socket, worth glancing at rather than guessing - and openpose_json, for nodes that consume keypoints directly instead of as an image. On a sequence it carries the first frame's, so the socket keeps its shape.
Installing it
Via ComfyUI Manager, search On-Set Studio. Or by hand:
cd ComfyUI/custom_nodes
git clone https://github.com/IsItDanOrAi/ComfyUI-On-Set-Studio
Restart ComfyUI. That's the whole install: the editor ships prebuilt, there's no Node.js step, nothing to compile, and the pack declares zero Python dependencies - its pyproject.toml is emphatic about that and the shipped code agrees. opencv-python and imageio are imported lazily, and only by Plate Cache when you hand it a video_path.
Two things you supply. A character - none ships, because Mixamo rigs are Adobe's to distribute. Grab X Bot or Y Bot from mixamo.com as an FBX and drop it into ComfyUI/custom_nodes/ComfyUI-On-Set-Studio/editor_dist/models/, or just import it in the app. And a GPU plus a browser that likes WebGL, since the editor is real-time 3D in a tab; the pack's Python is CPU-only.
Optional: ARDY, NVIDIA's text-to-motion model, as a separate local service on port 8765 wanting ~14GB of VRAM. Entirely skippable if you'd rather keyframe.
What goes wrong, and what's just expected
The editor looks broken and greyed out. It isn't. The Scene tab stays dead until you import an FBX and press the Off button at the top right of the Scene panel so it reads On. The ground grid and green screen are behind the same switch. The Environment tab works with no character at all, which is where people give up too early - and the 404 the console logs for a missing character model on a fresh install is expected, not a fault.
"no payload for session 'default'." The node has nothing to hand out. Open the editor, build something, press Send to ComfyUI, then Run. It won't invent a scene.
A tiny 64x64 black image came out of your graph. That's the placeholder for a slot that wasn't rendered. Check slot_labels.
A sequence came out short. Interrupted sends emit the frames that exist and warn you. Changing output resolution mid-send drops mismatched frames instead of crashing.
Only Mixamo-style bone naming is understood. Biped, Unreal Mannequin or Rokoko rigs need converting, and non-standard proportions make ARDY's retarget stride wrong until you regenerate mixamo_rest.json.
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| session | STRING | default | — |
Outputs (9)
| Name | Type | Description |
|---|---|---|
| out_1 | IMAGE | — |
| out_2 | IMAGE | — |
| out_3 | IMAGE | — |
| out_4 | IMAGE | — |
| character_json | STRING | — |
| environment_json | STRING | — |
| scene_json | STRING | — |
| slot_labels | STRING | — |
| openpose_json | STRING | — |