STG: Spatiotemporal Skip Guidance, attention skip (Hyung et al. 2025)
Skip the attention layer entirely and steer away from what's left
- model
- MODEL
The weak-branch family all share the same move: run the model again in a deliberately worse configuration and steer away from that prediction. PAG breaks attention into identity, SEG blurs it, and STG just... removes it.
STG is Spatiotemporal Skip Guidance (Hyung et al., CVPR 2025), originally about skipping layers in video diffusion. This node is the self-attention skip - the self-attention sublayer's contribution is dropped for the chosen blocks, on the theory that it adds little, so what remains is a blunter version of the same prediction. It's exact on Anima, whose attention carries no bias, and it works on SD 1.5 and SDXL too.
The mechanism
One extra model pass per step, with the selected self-attention layers skipped. Then:
out = mix + s * (c - skip)
s is the strength, and the default is 1.0 - "paper 1 for the residual skip" - which is a much smaller number than PAG's 3 or SEG's 3. That's not timidity: the skip is a large perturbation, so a small weight goes a long way. If you're used to dialling PAG up to 3, don't do that here.
Because the weak branch is a reduction rather than a replacement, it costs a full extra forward pass per step but nothing else - no extra state, no second prompt, and it composes with whatever your other stages are doing.
Inputs and output
model- loader → node → sampler.scale- default 1.0. "Strength s (paper 1 for the residual skip)". The useful range is roughly 0.5–2; past 3 you're steering away from something so degraded that prompt adherence starts to suffer.blocks- which self-attention blocks get skipped. Same preset list as SEG:middle (PAG / SEG default),middle + first output,deep output (output 0-2),deep input (input 7-8),all SDXL attention blocks. On Anima and other Cosmos-Predict2 transformers the labels map to the two middle blocks, the first third and the last third.
Output: MODEL.
When to pick STG over SEG or PAG
My honest reading of the trade-offs:
- On Anima, STG is a first-class citizen. Along with SEG it's one of the only weak branches that can run on a Cosmos-Predict2 transformer, and it's the cheapest conceptually - no blur radius to tune, one scale, one block choice.
- On SDXL, SEG is usually the better starting point. Blurring queries gives you a dial (that
blur_sigma) between "nearly plain" and "fully flattened"; skipping has no such dial, and the strength scale is the only control. Community reports on PAG in this family also point out how easy it is to overshoot on guidance-sensitive checkpoints - STG has the same failure mode, just at a different constant. - If you want to mix and match blocks or modes, use the stage node.
CFG Weak Branch: Perturbed Self-Attentionexposes all four methods (PAG, SEG, temperature, skip) plus a mode switch between "add on top of CFG" and "replace the unconditional". This node is the paper's version with the paper's defaults.
Chaining STG and SEG does nothing: the weak-branch stage holds one rule, and the later node replaces the earlier one. Same for the PAG and SEG paper nodes sitting in the same folder.
Install
Manager → search CFG Megapack → install, restart. Or:
cd ComfyUI/custom_nodes
git clone https://github.com/AbstractEyes/comfy-cfg-megapack
No dependencies, no downloads, no requirements.txt. The pack needs a reasonably current ComfyUI - it uses comfy_api.latest, tested on 0.38.0 - so on an old install the nodes won't be there at all after a restart.
Where people get burned
Turning the scale up. STG's default of 1 is the paper's value, and 1 is a lot. Doubling it does not double a subtle effect; it changes the image substantially, often washing out fine detail. Move in 0.25 steps.
Expecting a free speed-up. The skip is only on the second pass. You still pay one extra full forward per step, which on SDXL is roughly a 50% hit to sampling time. There's no mode here that saves you the unconditional pass - that trick lives in the stage node's "replace" mode.
Nothing happening. As always with this pack, the suspect is ComfyUI's single CFG-function slot: a stock RescaleCFG, Mahiro or RenormCFG node chained after this one takes it and your weak branch does nothing. CFG Plan Readout prints the plan stage by stage, so one queue tells you whether STG is installed.
Using it at the same time as a high cfg. The weak branch adds guidance on top of CFG. On a CFG-sensitive checkpoint, drop to cfg 4–7 when you enable it, then adjust the strength.
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| scale | FLOAT | 1.00–20 | Strength s (paper 1 for the residual skip). |
| blocks | COMBO | middle (PAG / SEG default) | Which self-attention blocks are perturbed (SDXL names; the middle block exists on SD1.5 too). On Anima and other Cosmos-Predict2 transformers: middle = the two middle blocks, deep input = the first third, deep output = the last third. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |