AdaMaG: adaptive manifold guidance (Esmati et al. 2026)
Guidance that fades out as the image appears
- model
- MODEL
AdaMaG (adaptive manifold guidance, Esmati, Hyung, Dadashzadeh, Choo & Mirmehdi, arXiv 2026) is a fresh-off-the-press method, and it's built on an observation you may have had yourself: guidance matters enormously at the start of a run and much less at the end. Early on, the latent is mush and the unconditional prediction is a genuine alternative reading of the scene; late, everything has settled and CFG mostly just adds contrast.
So AdaMaG does two things at once.
The two mechanisms
A decaying scale. The guidance scale at each step is max(w_min, w · t^gamma) - the scale you
asked for, multiplied by the noise level t raised to gamma, with a floor underneath it. At high
noise (t near 1) you get close to your full scale. As t falls, the scale falls with it, rapidly
if gamma is large. The default gamma is 4, which is aggressive: by the time t is 0.5 the scale
is down to a sixteenth of w, and normally you're already sitting on the w_min floor of 1 by the
middle of the run. The node's model page says exactly this - the scale falls off as the noise level
falls.
A partial projection. The guidance difference is split along the conditional noise estimate
into the part parallel to it and the rest. beta (default 0.1) is how much of the parallel part
survives. This is the APG/Tangential-Damping family of move - those methods drop the parallel
component because it's the one that raises saturation; AdaMaG keeps a sliver of it instead of
discarding it outright.
Put together: strong, mostly-projected guidance while the image is still being decided, then almost none. If you've ever wished you could get the composition of CFG 7 and the finish of CFG 3, that's the shape of the intent.
Inputs
- model - from the loader, before the sampler.
- scale (default -1) -
-1uses the KSampler's cfg aswin the decay formula. Because the schedule eats most of it, people run this at higher base scales than they would with plain CFG. - beta (default 0.1) - how much of the along-the-noise-estimate part to keep.
0is a full projection (the parallel component is discarded);1is much closer to plain CFG. - gamma (default 4) - how fast the scale decays. This is the input that actually changes the feel of the node. Lower it (1–2) if the image is losing prompt adherence in the late steps, raise it if you want the tail of the run almost entirely unguided.
- w_min (default 1) - the floor.
1means the late steps fall back to the conditional prediction alone; raise it slightly (1.5–2) if the finish looks under-guided. - space -
autouses the method's published space. Nonlinear rule, so the space is a real decision; leave it onautounless you're deliberately experimenting.
The output is a MODEL, dropped between the loader and the sampler like every other patch node in the pack.
Install
# ComfyUI Manager: search "CFG Megapack" -> Install -> restart
# or:
comfy node install comfy-cfg-megapack
# or by hand:
cd ComfyUI/custom_nodes && git clone https://github.com/AbstractEyes/comfy-cfg-megapack
The pack ships no requirements.txt because it needs nothing beyond torch and the standard library,
and there are no model files to download. ComfyUI 0.38 or newer is required.
Honest caveats
This is a 2026 paper with a handful of citations and, at the time of writing, no community body of
practice around it. There is no "here's what everyone runs" setting. Treat the defaults as a starting
point and use the pack's own instrumentation: CFG Measure: Per-Step Probe logs the scale actually
used at each step to output/cfg_probe/, which is how you confirm the decay curve is doing what the
formula says. CFG Plan Readout prints the installed plan if you've chained several nodes and lost
track of which one is in charge.
Two structural notes. AdaMaG writes the combine stage of the pack's seven-stage guidance plan, so a
later combine node replaces it entirely; corrections, by contrast, stack. And the pack's hook forces
the unconditional pass on every step it guides, so the CFG-1 speedup you might be relying on for a
distilled model goes away as soon as any of these nodes is on the model wire - which for AdaMaG is
also moot, since at w = 1 there's nothing to project.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| scale | FLOAT | -1.0-1–100 | The guidance scale w for this rule. -1 uses the sampler's cfg value. |
| beta | FLOAT | 0.100–1 | Weight of the part along the noise estimate. |
| gamma | FLOAT | 4.00–10 | How fast the scale falls with the noise level. |
| w_min | FLOAT | 1.00–10 | Lowest scale. |
| space | COMBO | auto (the method's own) | Where the rule is computed. Linear rules give the same image in any space; nonlinear ones do not. 'auto' uses the space the method was published in (noise for most, denoised for APG and the angle rule, velocity for flow models). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |