Parallel Upscale (Multi-GPU)
One node, every card working at once
- upscale_model
- image
- start_a
- start_b
- start_c
- start_d
- start_e
- start_f
- start_g
- start_h
- IMAGE
This is the "just make it faster" node. Instead of the pack's split-into-branches-and-merge choreography, Parallel Upscale (Multi-GPU) takes an upscale model, an IMAGE batch, and a comma-separated list of GPU ids - and when you hit Queue it spreads the batch across those cards and upscales everything at once. One node, no Start plumbing, no separate merge step.
How it works
You give it gpu_ids like 0,1,2,3. The node parses that into a list of devices, then round-robins frames across them: frame 0 to card 0, frame 1 to card 1, frame 2 to card 2, frame 3 to card 3, frame 4 back to card 0, and so on. Each worker clones the upscale model (so cards never fight over one shared nn.Module), runs the same tiled upscale the stock ComfyUI node uses, and the results are stitched back together in original order. Round-robin instead of contiguous chunks is a nice touch - if your cards aren't identical, each one gets a fair spread of the work instead of the slow card getting a contiguous tail.
The optional start_a … start_h inputs are just ordering hints: they're START sockets you can feed from GPU Init nodes, and the node touches them so ComfyUI is forced to run the Init nodes first. You can ignore them entirely if you don't care - the GPU context gets created either way on the first touch.
Inputs and outputs
- upscale_model - required, from the stock UpscaleModelLoader. Any ESRGAN/UltraSharp/RealESRGAN in
models/upscale_models/. - image - required, an IMAGE batch (a single image, many images, or video frames).
- gpu_ids - required STRING, default
"0", comma-separated device indices like0,1,2,3. The tooltip lists the cards ComfyUI can actually see. Leave it at the default and you're just doing single-card upscale, which is fine, but not why you installed this pack.
Output: IMAGE, the full batch, upscaled and back in original order. It's the same shape as the input batch - just bigger pixels.
When to use it instead of the Split/Merge layout
For images, this is the one I'd reach for. It's a quarter of the nodes and does the same job. For video, the Split Image Batch → per-GPU branches → Merge Image Batch layout is the better fit - it's the pack's documented video recipe, it lets each branch's async overlap kick in, and you can see each card's work as separate branches instead of one opaque blob.
The trade-off with this node: it's a monolith. One card OOMs and the whole node errors, and you can't watch per-card progress. If you're chasing a VRAM issue across four cards, the explicit branch layout is much easier to debug.
Install and gotchas
# ComfyUI Manager → Install Custom Nodes → search "32GPU Video Upscale"
# or:
cd ComfyUI/custom_nodes
git clone https://github.com/WhyNotNN/ComfyUI-32GPU-Video-Upscale.git
# restart ComfyUI
No extra pip dependencies. Two gotchas worth remembering: each listed GPU must actually be visible (asking for card 3 with two cards installed is a hard error), and each worker loads its own copy of the model, so VRAM use is per-card - a 4x upscaler on each of your 24GB cards is no problem, but don't expect the batch to be split across one card's memory. On a CPU-only box it falls back to CPU and becomes a slow way to do nothing useful. And the usual caveat: this is a brand-new pack with no real community track record yet, so treat early behavior with a skeptical eye.
Inputs (11)
| Name | Type | Default | Description |
|---|---|---|---|
| upscale_model | UPSCALE_MODEL | — | |
| image | IMAGE | — | |
| gpu_ids | STRING | 0 | Comma-separated GPU indices, e.g. 0,1,2,3 CUDA/ROCm not available (CPU fallback) |
| start_aopt | START | — | |
| start_bopt | START | — | |
| start_copt | START | — | |
| start_dopt | START | — | |
| start_eopt | START | — | |
| start_fopt | START | — | |
| start_gopt | START | — | |
| start_hopt | START | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| IMAGE | IMAGE | — |