Inpainting VAE Encode
Grow the mask before the sampler sees it
- pixels
- vae
- mask
- LATENT
Inpainting in ComfyUI comes down to one trick: hand the sampler a latent that says which pixels are fair game. Outside the mask, the sampler keeps what's already there; inside, it denoises. Everything else - denoise strength, model choice, prompt - is downstream of getting that mask boundary right.
The boundary is usually wrong by a few pixels, and that's the gap this node fills. Core's VAE Encode (for Inpainting) takes your mask exactly as drawn. This one lets you push it outward or pull it inward first.
How it works
Four steps, and knowing them explains the surprises.
The mask gets stretched to the image's size first, so a mask drawn at another resolution still lines up. Then the image is cropped to the nearest multiple of 8 - a latent addresses pixels 8 at a time, so the trim is taken evenly off both sides and your output can be a few pixels smaller than your input. That's the "why is my 1024 image now 1016" question, and it isn't a bug.
Then mask_offset does its work. Positive grows the white area by convolution, negative shrinks it, 0 leaves it alone. The mask is rounded to 0 or 1 before that, so grey edges become hard edges - you get a clean grow/shrink, not a feathered one.
Last, the masked pixels are flattened to mid grey (0.5) before encoding. That detail matters more than it looks: an encoder reads a black patch as content - a dark shape you asked for - while flat mid grey reads as absence.
Output is a LATENT whose noise_mask is the adjusted mask you actually built, not the one you drew. Wire it into a KSampler, where the noise mask is what limits the denoising to the region, and decode as usual.
mask_offset is the whole point
Default is 6, and it's the number to play with. Positive, around 6 to 12, grows the painted area past what you drew - this is what hides the seam where new content meets old, and it's the right call for removing an object, because the old object's shadow and edge pixels live outside your hand-drawn mask. Negative shrinks it, keeping more of the original: use it for touching up the middle of a face or a garment without letting the sampler wander into the outline you liked.
Range is −128 to 128. If you find yourself past 20 in either direction, the mask was probably the problem.
Install
ComfyUI Manager, search WAS Node Suite v3, install, restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/WASasquatch/was-node-suite-comfyui.git
No packages, no downloads, no build step - ComfyUI 0.14.0+ and Python 3.10+. First start writes config.yaml and the pack's folders, plus a bytecode compile, then starts are normal. The node is in the extras group, on by default. Old guides telling you to install opencv for this pack are from v2, which pinned packages and broke on ComfyUI updates; v3 installs nothing.
Where people get burned
Colours shift in the whole image, not just the masked part. This is the oldest complaint about inpainting and it isn't this node's fault: the entire image goes through VAE encode and decode, and that round trip moves colours slightly. Two ways out - composite only the masked pixels back over the original with a mask composite when you're done, or use a crop-and-stitch inpaint setup that never sends the unmasked pixels through the VAE at all. The crop/stitch pair is the more structural answer and it's what most people have moved to.
The VAE has to match. vae should be the one belonging to the checkpoint the sampler runs - load it from the same source, not a separate VAE from a different era. Mismatched VAEs produce colour and structure drift that gets blamed on the sampler.
A hard seam around a removed object. Your mask_offset is too small, or the mask was drawn tightly to the object's visible edge rather than past it.
Nothing gets repainted. Check the sampler's denoise. A latent with a noise mask still needs enough denoise to actually change the region; at very low values you get back roughly what you put in and conclude the node is broken.
A grey or shaded box you didn't ask for. If you're running a face pass after upscaling, the encode/decode cycle is the usual culprit, and it compounds with each pass. Increase padding, verify the VAE, and don't stack more VAE round trips than you need.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| pixels | IMAGE | The image to inpaint. Its width and height are cropped to the nearest multiple of 8, taking the trim evenly from both sides. A latent addresses the image 8 pixels at a time. | |
| vae | VAE | The VAE that turns the prepared image into a latent. Use the one that belongs to the checkpoint the sampler runs, or the colours shift. | |
| mask | MASK | Which part is repainted. White is repainted, black is kept, and grey is rounded to one or the other. It is stretched to the image's size first, so a mask drawn at another resolution still lines up. | |
| mask_offset | INT | 6-128–128 | How far the painted area grows or shrinks before encoding, in pixels of the input image. 0 uses the mask exactly as drawn. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| LATENT | LATENT | The encoded latent with the adjusted mask attached as its noise mask. Feed it to a KSampler, which will only replace the masked part. |