VAGS: velocity-adaptive guidance scale (Luo et al. 2026)
Let the two predictions vote on how much guidance you need this step
- model
- MODEL
Most adaptive guidance methods are one-note: they notice one signal and modulate the scale with it. VAGS (Luo et al., arXiv 2026) watches two at once - how far along the run you are, and how much the two predictions agree - and multiplies your scale by a factor derived from both:
w_i = w * exp(kappa * (2s - 1) * cos(u, c))
s is the signal level (1 at pure noise, 0 at the clean end), and cos(u, c) is the cosine between the unconditional and conditional predictions. Early in the run (2s - 1) is positive, so agreement produces more guidance; near the end it flips negative, so agreement produces less. The paper's reading is that a step where the two predictions point the same way late in the run is a step where extra guidance mostly amplifies what's already there.
One multiplication, one cosine, no extra forward pass. The pack measures it at about 1–2% wall clock, which is the cheapest adaptive method in the folder.
Inputs and output
model- loader → node → sampler.scale- the basew; -1 takes the sampler's cfg.kappa- default 1.0, range 0–5. "Strength of the adaptation (0 = plain CFG)." The paper used 0.9 on Flickr30K, so 1.0 is already at the upper end of the published range; treat 2+ as an experiment.space-auto (the method's own), which on a flow-matching model means the velocity.
Output: MODEL.
Because kappa 0 is exactly plain CFG, this node is trivially A/B-able: same seed, kappa 0 versus 1, nothing else changed. That's the first thing I'd do with it.
What to expect, honestly
On an eps-prediction SDXL-family checkpoint the cosine between u and c sits extremely close to 1 for most of the run - the two noise predictions differ by about a percent of their size - so the multiplicative factor is nearly constant and the whole adaptation is gentle. You'll get a subtle redistribution: slightly more guidance early, slightly less late.
On a flow-matching model the cosine moves more, and there's more for the rule to work with. That's where the method was developed, and where it's worth trying.
If you want a visible adaptive-scale effect on SDXL, the schedules (CFG When, TV-CFG, Wang's) will do more for less fuss. VAGS is the elegant version of a small idea; it's not a rescue for a bad prompt.
Install
Manager → search CFG Megapack → install → restart. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/AbstractEyes/comfy-cfg-megapack
Nothing to fetch - no model files, no requirements.txt, no API key. The pack's one hard constraint is ComfyUI freshness: it's written against comfy_api.latest and the README reports testing on 0.38.0 with torch 2.11, CUDA 12.8 and CPU-only. Old builds won't import the pack at all.
Where people get burned
Turning kappa up to "make it stronger". The factor is exponential in kappa, so 3 or 4 doesn't mean "three times the effect", it means the scale swings by multiples across the run - which usually shows up as some steps over-guided and others barely guided at all. Stay near 1.
Reading its effect as a quality jump. This is a fine-tuning method. If your images are bad at cfg 7, VAGS at kappa 1 will not fix them; it will produce a slightly differently-shaped version of the same image.
Assuming it stacks with a schedule. It writes the combine stage, so it does stack with CFG When and the schedule paper nodes - the schedule sets the scale, VAGS modulates it. That's a legitimate combination, but it makes attribution hard. Change one thing at a time, and use CFG Measure: Per-Step Probe if you want to see the per-step scale that actually ran rather than the one you typed.
The single CFG slot. A stock RescaleCFG, Mahiro or RenormCFG chained after this node takes ComfyUI's CFG-function slot and VAGS goes quiet. CFG Plan Readout will show you the plan if the image doesn't budge.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| scale | FLOAT | -1.0-1–100 | The guidance scale w for this rule. -1 uses the sampler's cfg value. |
| kappa | FLOAT | 1.000–5 | Strength of the adaptation (0 = plain CFG). |
| space | COMBO | auto (the method's own) | Where the rule is computed. Linear rules give the same image in any space; nonlinear ones do not. 'auto' uses the space the method was published in (noise for most, denoised for APG and the angle rule, velocity for flow models). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |