Contextual Diffusion Options
Stop your tiles from inventing a second cathedral
- options
- options
Tiled sampling has a failure mode nobody warns you about, and it's the reason this node exists. The author's description of hitting it is the most honest thing in the pack's README: editing a 2160×3072 image with FLUX.2 Klein 4B, ordinary tiled diffusion looked great until he zoomed out. "One tile had found a figure, another had invented a second figure, and different parts of the cathedral had become different buildings. The overlaps were smooth! The scene was still nonsense."
That's the whole problem in one paragraph. Tiling keeps each evaluation small; it does not let any tile see the complete image. Contextual Diffusion fixes that by making the sampler look at the whole picture and full-resolution detail during the same denoising pass.
How it works
For each denoising step you get two predictions: the normal set of local contexts, plus one prediction of a smaller, aspect-preserving view of the entire image. The whole-image branch is what keeps the composition coherent. Directly blending it in makes everything blurry, so it's added as a correction instead - the global prediction minus the low-frequency content the local tiles already produced:
prediction = local + scheduled_weight × (global_upsampled − local_low_frequency)
The global voice starts strong and fades. global_steps is how many initial steps get the correction at all, and global_decay multiplies that strength after each one - with the defaults (global_steps: 1, global_decay: 0.5) you're correcting the first step at full weight and nothing after. On a distilled model like Klein that finishes in four steps, one corrected step is already a quarter of the whole run, so it's less of a dial than it looks.
Credit where it's due: the author says he developed this independently and only later found the low-frequency/residual split resembles Upsample Guidance (arXiv 2404.01709), and he's explicit that his evidence is qualitative. This is his method, not a standard.
The inputs that matter
latent_context_size- default 96, max 512. This is the size of the local window and it's square: both local dimensions come from this one number, in latent pixels.global_steps/global_decay- the two you'll actually turn. Raiseglobal_stepsif the scene drifts apart, lower it (or lowerglobal_decay) if the result looks smoothed and low-frequency garbage is creeping into the final detail.global_weight- default 1, range 0 to 2. How authoritative the whole-image branch is. Lower it to let the tiles interpret more.latent_context_overlap(default 32) andlatent_context_batch_size(default 4) - same tradeoffs as any tiled sampler: overlap costs work, bigger batches cost memory.diffusion_mode-multidiffusionormixture_of_diffusers, controlling how overlapping local contexts blend.
The output is a single options socket that goes into another options node or the KSampler. Plan for the cost: one tiled pass every step, plus one smaller whole-image evaluation for each step that uses the correction.
Install
Same pack, same steps:
cd ComfyUI/custom_nodes
git clone https://github.com/Artificial-Sweetener/SimpleSyrup.git
cd SimpleSyrup
python -m pip install -r requirements.txt
Then restart. Or use Manager: search the Node Pack list for SimpleSyrup, install, restart. You need a current ComfyUI - the pack uses the v3 extension API, so an old install just silently has no nodes.
When it's the wrong tool, and the minefield
This is for edits the model already knows how to make at a normal resolution - clothing, materials, color, jewelry, expression, local lighting - where you uploaded a big source and want to keep the pose and composition. If you need a new pose, camera and environment, get those at normal working resolution first and refine afterward. No amount of whole-image guidance invents a composition the model would never have produced at that canvas size.
Then the constraints, all from the sampler's own validation:
- Connect Tiling Options too and yours is ignored. Contextual Diffusion takes precedence, logs a warning, and substitutes its own local plan.
- UniPC is out.
uni_pc/uni_pc_bh2raise "Tiling and Contextual Diffusion do not support UniPC samplers." - No ControlNet, no regional conditioning, no GLIGEN. Conditioning with
control,areaorgligenkeys is rejected - the author hasn't validated their spatial behavior across two context sizes and says so plainly. latent_context_sizemust be at least 16 and overlap has to be smaller than the local window, or the node throws before a single step runs.
One you'll actually use: the KSampler's segs input swaps the regular grid for region-guided windows, so detected objects land inside their own contexts instead of being sliced down a tile boundary. An earlier version ran a second bank of SAM views alongside the tiles and produced duplicate hats and extra limbs; the author removed that. One plan, and the dedicated KSampler (Contextual Diffusion) returns the windows it used as contexts_segs so you can see what it evaluated.
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| latent_context_size | INT | 9616–512 | Maximum side of each model context in latent pixels. Larger contexts preserve more relationships but use more memory. |
| global_weight | FLOAT | 1.000–2 | Strength of whole-image low-frequency guidance. 1 makes the global context authoritative; lower values allow more tile interpretation. |
| global_steps | INT | 10–10000 | Number of initial denoising steps that use the global context. Fewer steps leave more late sampling for local detail. |
| global_decay | FLOAT | 0.500–1 | Multiplier applied to whole-image strength after each global step. Lower values hand control to local contexts faster. |
| diffusion_mode | COMBO | multidiffusion | Tile overlap blend. MultiDiffusion averages predictions; Mixture of Diffusers gives tile centers more influence. |
| latent_context_overlap | INT | 320–256 | Overlap between local latent contexts in latent pixels. Larger overlaps reduce seams but increase sampling work. |
| latent_context_batch_size | INT | 41–8 | Number of equal-sized latent contexts sampled together. Higher values can be faster but use more memory. |
| differential_diffusion | BOOLEAN | false | Uses the noise mask to vary denoising strength spatially; preserves existing model mask behavior. |
| optionsopt | SIMPLE_SYRUP_SAMPLER_OPTIONS | Optional preceding sampler options; bypass this node to omit its contribution. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| options | SIMPLE_SYRUP_SAMPLER_OPTIONS | Combined sampler options; connect another options node or KSampler. |