Nodes/was-node-suite-comfyui/Latent Hybrid Upscale
ComfyUI Node Runs on cloud

Latent Hybrid Upscale

Smooth the edges, leave the flat bits alone

By WASasquatch·Created 4 years ago·Updated a day ago· 1,864
Latent Hybrid Upscale
  • latent
  • vae
  • donor_latent
  • latent
  • edge_mask
◄scale2.00►
◄pre_blur_sigma_px1.00►
◄canny_threshold125►
◄canny_threshold2155►
◄canny_l2gradienttrue►
◄dilate_radius_px8►
◄feather_sigma_px6.00►
◄mask_min0.00►
◄mask_max1.00►
◄use_nearest_exacttrue►
◄output_mask_resolutionimage►
◄video_decode_horizontal_tiles2►
◄video_decode_vertical_tiles2►
◄video_decode_overlap_latent4►
◄video_decode_last_frame_fixfalse►
◄video_decode_enable_cudnntrue►

Every "hires fix" graph needs one boring step: make the latent bigger before you sample it again. Core ComfyUI's latent upscale does it by interpolation - bilinear or nearest - and both are wrong in a visible way. Bilinear softens the whole thing, nearest leaves you looking at stair-stepped blocks. Latent Hybrid Upscale refuses to choose. It decodes the latent, finds the outlines in the picture, and only goes smooth there.

What it actually does

Three things happen under the hood, in this order.

First the node builds two enlarged versions of your latent: a crisp one, where each output block copies the nearest source block (use_nearest_exact picks the centre-corrected version of that rule), and a smooth bilinear one. Then it decodes the original latent through the vae you wire in purely to look at it, finds edges with a Canny pass, dilates and feathers them into a mask, and blends: white on the mask, you get the smooth bilinear enlargement; black, you keep the crisp block copy.

That's the trade it's making. Flat areas - skin, sky, walls - stay exactly as sharp as the latent had them, no bilinear mush. Outlines get interpolated, so no stair-stepping along a jawline. Since flat regions are most of a typical image, this usually reads as "sharper than a plain latent upscale" rather than "smoother", which is not what the name suggests.

Outputs are latent (wire it into a KSampler at a low denoise, or straight into VAE Decode if you just want the bigger picture) and edge_mask (a MASK, at picture or latent resolution depending on output_mask_resolution). Watch that mask while tuning. If the mask is white everywhere, you've built an expensive bilinear upscale; if it's black everywhere, you built a nearest one. It's the honest readout for the Canny settings.

The inputs you actually touch

  • latent, vae and scale. Those three get you running. Scale rounds to whole latent blocks, so 1.5 on a 128-wide latent lands on a size you didn't quite ask for.
  • canny_threshold1 / canny_threshold2 (defaults 25 / 155) with pre_blur_sigma_px (1.0) in front of them. The blur is the important one: it stops film grain and fabric weave from registering as edges. Push it to 2.0 and above and only major outlines survive.
  • dilate_radius_px (8) and feather_sigma_px (6). Dilate is what gives the smooth pass a band wide enough to hide the stair-step, feather hides the band's own border. Leave them.
  • mask_max is your escape hatch. Drop it below 1.0 and you never fully commit to the smooth version, which is what you reach for when outlines come out soft.

On a video latent the decode for edge-finding is split across video_decode_horizontal_tiles × video_decode_vertical_tiles, cross-faded by video_decode_overlap_latent, so a long sequence doesn't need the VRAM of one full-frame decode. The optional donor_latent takes the smooth half from a second latent instead of interpolating yours - a way to combine two latents, not a beginner move.

Install

ComfyUI Manager, search for WAS Node Suite v3, install, restart. Manually:

cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git

That's it. The pack installs no packages and downloads nothing - v2's list of twenty pip dependencies is gone in v3 (a ComfyUI-driven opencv upgrade fighting insightface was the classic way this pack broke after an update; there's nothing left to fight over). Needs ComfyUI 0.14.0+ and Python 3.10+. First start writes config.yaml and the pack's own folders under your ComfyUI user dir and takes a second longer than later starts.

Latent Hybrid Upscale sits in the extras feature group, which is on by default. If you also run the standalone WAS Extras pack, that group gets switched off in config.yaml - which is why the node can vanish from the menu after someone edits that file.

Where people get burned

The VAE has to match the latent. It only reads, but the edges it finds are the whole point. Feed it a latent encoded with a different VAE and the outlines land in the wrong places - the mask looks plausible and the result is subtly smeared in odd spots. Empty the vae input and the node tells you rather than guessing.

The decode costs time and VRAM. This is the entire expense of the node. On a big latent or a video sequence, that decode is what blows up your card, not the upscale - cut it into tiles instead.

Pre-blur at 0.0 finds everything. Grain registers as edges, the mask covers the frame, and you've built bilinear upscaling with extra steps.

It's still latent space. This doesn't invent detail, it just gives the second sampler a cleaner canvas to add detail onto. If your flat areas look hard-edged and blocky afterwards, that's the crisp half doing its job and your denoise is too low - not the node failing.

CategoryWAS Suite/Latent/Transform

Inputs (19)

NameTypeDefaultDescription
latentLATENTThe latent to enlarge. A video latent with a time axis is handled frame by frame.
vaeVAEThe VAE used to decode the latent so its edges can be found. It must be the one that matches the latent, or the edges will be found in the wrong places. It only reads the latent; the result is still built in latent space.
scaleFLOAT2.001–8How much larger the result is. 2.0 doubles both sides, 1.5 adds half again. Sizes are rounded to whole latent blocks.
pre_blur_sigma_pxFLOAT1.000–20Blur applied to the decoded picture before edges are looked for, in pixels. It keeps film grain and fine texture from registering as edges; 0.0 finds every last one, 2.0 or more keeps only the major outlines.
canny_threshold1INT250–1000The lower of the two edge-detection levels, on a 0-255 scale. A faint edge is kept only when it joins a strong one, and this is how faint it may be. Lower values trace more of an outline; raise it if speckles appear in flat areas.
canny_threshold2INT1550–1000The upper of the two edge-detection levels, on a 0-255 scale. Anything this strong starts an edge on its own. Raise it to keep only bold outlines, lower it to catch soft ones.
canny_l2gradientBOOLEANtrueHow edge strength is measured. On, the true length of the gradient is used, which is slightly slower and more accurate on diagonals. Off, a cheaper approximation is used that reads diagonal edges as stronger than they are.
dilate_radius_pxINT80–64How far the found edges are grown, in pixels. Edges are hairline by themselves, so growing them is what gives the smooth enlargement a band to work in: 8 covers a typical outline, 0 leaves the raw one-pixel lines.
feather_sigma_pxFLOAT6.000–50How far the grown edge fades out, in pixels. Without it the band would have a visible border of its own; 6.0 gives a soft changeover, 0.0 leaves a hard-edged band.
mask_minFLOAT0.000–1Floor under the finished mask. Raise it above 0.0 to let a little of the smooth enlargement into areas with no edges at all, which takes the hard blockiness off the whole picture.
mask_maxFLOAT1.000–1Ceiling over the finished mask. Lower it below 1.0 to keep some of the crisp enlargement even on the strongest edges, which is the way back when outlines come out too soft.
use_nearest_exactBOOLEANtrueHow the crisp half of the blend is enlarged. On, each output block takes the value of the source block whose centre is nearest, which keeps the picture from drifting half a block sideways. Off uses the older nearest-neighbour rule.
output_mask_resolutionCOMBOimageWhich size the mask output comes out at. `image` gives it at the size the enlarged latent decodes to, ready to view or reuse against the finished picture. `latent` gives the small version that actually drove the blend.
video_decode_horizontal_tilesINT21–8How many columns a video latent is split into for the decode that finds edges. More tiles means less VRAM and more time. Ignored on an image latent.
video_decode_vertical_tilesINT21–8How many rows a video latent is split into for that decode. 2 rows and 2 columns is four tiles, each a quarter of the frame. Ignored on an image latent.
video_decode_overlap_latentINT40–32How far neighbouring tiles overlap, in latent units. The overlap is cross-faded, so raise it if seams show along the tile boundaries; 0 turns the fade off entirely.
video_decode_last_frame_fixBOOLEANfalseWhether the final frame is duplicated before decoding and the extra output dropped afterwards. Turn it on when the last frames of a clip decode to something corrupt, which some video VAEs do.
video_decode_enable_cudnnBOOLEANtrueWhether cuDNN is left on for the video decode. Turning it off is slower and avoids the large workspace allocations that make some cards run out of memory part way through a clip.
donor_latentoptLATENTWhere the smooth half of the blend comes from. Leave it unconnected and the node interpolates the input latent. Connect a second latent of the same batch and channel shape, a version sampled at a higher resolution, say, and its detail is what gets laid into the edges.

Outputs (2)

NameTypeDescription
latentLATENTThe enlarged latent, ready for a sampler or a decode.
edge_maskMASKWhere the smooth enlargement was used: white along the edges the node found, black elsewhere. Watch it while setting the two Canny levels, or reuse it to treat the same areas downstream.