ControlNet Stack πΉοΈποΈ
Three ControlNets, One Node β Stacking Pose, Depth and Canny Without the Spaghetti
- positive
- negative
- control_net_1
- image_1
- control_net_2
- image_2
- control_net_3
- image_3
- control_net_4
- image_4
- vae
- positive
- negative
- summary
Stacking two ControlNets used to mean two ControlNet Apply (Advanced) nodes with the conditioning threaded through both, and every time you wanted to swap the depth map for a canny pass you were rewiring three wires and hoping you didn't cross them. OmniNodes' ControlNet Stack collapses that into one node with up to four slots. That's the whole pitch, and for this specific job it's a good one.
Why you'd stack at all
ControlNet is still the single most useful trick after text-to-image itself: the prompt decides what appears, the condition decides where it goes - edges, depth, pose, normals. And the realistic 2026 workflow is two or three conditions at once, because each one is bad at what the others are good at. Canny nails hard architectural edges and fights you on anything organic. Depth gets the scene laid out correctly and doesn't care about texture. OpenPose pins a body and says nothing about the room.
Worth knowing before you build a four-slot monster: the modern union checkpoints ship fewer conditions than the SDXL era did, and they publish lower control weights than the old 1.0 default. Canny, depth, pose, scribble and gray are the menu on a 2026 union; segmentation, normal maps and QR-code brightness exist only on SD 1.5 and SDXL. Also, if you're on an SDXL union, its fusion is good enough that stacking two conditions usually wants less strength per slot, not more.
How it actually works
No magic here, and that's the point. The node imports ComfyUI's own core ControlNetApplyAdvanced and calls it once per active slot, feeding each result into the next. Slot 1 applies, its output conditioning becomes the input to slot 2, and so on through slot 4. The effect is identical to chaining that many Apply nodes by hand - the stacking order is the node's only invention. Nothing gets merged, no model weights are touched; you're just appending control entries to the same conditioning pair.
A slot only fires if both its control_net and its image inputs are connected. Leave a ControlNet input unwired and the slot is skipped entirely, widgets and all. Set its strength to 0 to disable a slot you want to keep plugged in.
Inputs and outputs
Required: positive and negative (CONDITIONING), plus the whole of slot 1 - control_net_1, image_1, strength_1, start_percent_1, end_percent_1. Slots 2 to 4 are optional and identical in shape. There's also an optional shared vae, only needed if one of your ControlNets expects a latent-space hint.
Three fields are the ones you actually turn:
strength_x- defaults to 1.0, range 0β10. On a modern union, 0.65β0.8 per slot is a saner start; 1.0 on two slots at once is how you get an over-cooked composition that ignores your prompt.start_percent_x/end_percent_x- when in the denoise schedule that condition applies. The standing advice for structure-heavy work is to release the condition once composition has formed: 0 β 0.5 for a canny or pose pass, letting the model add its own detail in the back half.- The summary output - a plain STRING listing every slot that fired with its strength and window, or telling you
No ControlNet slots active. Wire it into a text preview node. It's the difference between debugging a stack and guessing.
Outputs are positive and negative (both CONDITIONING, wire them into your KSampler) and that summary.
Installing it
Anything in the Model Utilities category runs on ComfyUI's own internals - no extra dependencies for this node.
# via ComfyUI Manager: search "OmniNodes", install, restart
# or manually:
cd ComfyUI/custom_nodes/
git clone https://github.com/TensorVizion/OmniNodes
Restart ComfyUI completely. You'll find it under TensorVizion/Model Utilities. Watch the terminal for [OmniNodes] β
Loaded lines; that's how the pack's __init__.py reports each file it registered.
Where people get burned
The pack is new (September 2026 additions) and has essentially no forum footprint - no wiki, no threads to search when something misbehaves. Read the summary output and the terminal, because that's the documentation.
Two I'd flag specifically. First, stacking costs VRAM for real: three SDXL ControlNets resident at once is three sets of weights, and that's the reason a plain T2I-Adapter still exists for people short on memory. Second, the depth pass in this same pack's ControlNet Preprocessor uses a depth_lite mode that estimates near-vs-far from luminance and local sharpness - fine for quick iteration, not comparable to a trained depth model. If your stacked depth condition isn't landing, that preprocessor is the first suspect, not the stack.
Inputs (23)
| Name | Type | Default | Description |
|---|---|---|---|
| positive | CONDITIONING | β | |
| negative | CONDITIONING | β | |
| control_net_1 | CONTROL_NET | β | |
| image_1 | IMAGE | β | |
| strength_1 | FLOAT | 1.000β10 | β |
| start_percent_1 | FLOAT | 0.0000β1 | β |
| end_percent_1 | FLOAT | 1.0000β1 | β |
| control_net_2opt | CONTROL_NET | β | |
| image_2opt | IMAGE | β | |
| strength_2opt | FLOAT | 1.000β10 | β |
| start_percent_2opt | FLOAT | 0.0000β1 | β |
| end_percent_2opt | FLOAT | 1.0000β1 | β |
| control_net_3opt | CONTROL_NET | β | |
| image_3opt | IMAGE | β | |
| strength_3opt | FLOAT | 1.000β10 | β |
| start_percent_3opt | FLOAT | 0.0000β1 | β |
| end_percent_3opt | FLOAT | 1.0000β1 | β |
| control_net_4opt | CONTROL_NET | β | |
| image_4opt | IMAGE | β | |
| strength_4opt | FLOAT | 1.000β10 | β |
| start_percent_4opt | FLOAT | 0.0000β1 | β |
| end_percent_4opt | FLOAT | 1.0000β1 | β |
| vaeopt | VAE | β |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| positive | CONDITIONING | β |
| negative | CONDITIONING | β |
| summary | STRING | β |