TSR: Temporal Score Rescaling (Xu et al. 2025)
Rescale the noise estimate so the sampler stops chasing the tail
- model
- MODEL
Most guidance nodes are built for SDXL, where cfg 7 is the middle of the road. TSR comes from the other end of the world - Temporal Score Rescaling (Xu, Wu, Park, Zhou & Tulsiani, ICML 2026) was developed on SD3 and FLUX, where guidance is already baked in and the interesting question is different: how do you get sharper images without turning the guidance back on?
The answer is to rescale the score by a signal-to-noise-dependent factor after CFG. Where the sample is noisy, leave the estimate alone. Where the sample is mostly decided, boost it toward the modes the model is confident about. That's a sharpening operation dressed up as guidance math.
The mechanism
Plain CFG first, then the noise estimate is multiplied by
r = (snr * tsr_sigma² + 1) / (snr * tsr_sigma² / k + 1)
with snr the signal-to-noise ratio at the current step. With k below 1 the numerator outgrows the denominator as the run progresses, so the rescaling starts neutral and grows sharper toward the clean end. With k = 1 the expression collapses to 1 and the node is off - hence the tooltip "1 = off".
The second knob, tsr_sigma, is the rescaling's own sigma, i.e. the point on the schedule where the transition happens. The paper used 3 on SD3/FLUX with k 0.93; the node ships k at 0.95 and tsr_sigma at 1.
No extra forward pass, no state, no buffer. It's arithmetic on a tensor you already have.
Inputs and output
model- the usual loader → node → sampler position.scale- thewthis rule uses; -1 means the sampler's cfg.k- default 0.95, range 0.5–1. "1 = off (paper SD3 / FLUX 0.93)". Lower means stronger mode-seeking.tsr_sigma- default 1.0, range 0.1–10. "The rescaling's own sigma (paper SD3 / FLUX 3)." On a flow-matching model you probably want to move this toward the paper's value; on SDXL's noise schedule the default keeps things conservative.space-auto (the method's own).
Output: MODEL.
Who should care
On a guidance-distilled model run at cfg 1 - Z-Image Turbo, Klein distilled, ERNIE Turbo - you have no guidance dial to turn. TSR is one of the few honest ways to sharpen output without pretending you're on SDXL: it operates on the prediction rather than on a difference between two predictions, so it doesn't need a working unconditional pass at all.
On SDXL it works but the payoff is smaller, and the defaults are deliberately mild because the paper's own tsr_sigma of 3 assumes a different schedule. If you try it on SDXL, start at the defaults and move k down in 0.01 steps.
Install
Manager → search CFG Megapack → install, restart. Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/AbstractEyes/comfy-cfg-megapack
Nothing to download, nothing to pip install, no requirements.txt. The one requirement is a recent ComfyUI - the pack is written on comfy_api.latest and tested on 0.38.0 - so if the nodes are missing after a restart, that's your first check, not your last.
Where people get burned
Pushing k down too far. Below about 0.8 the effect stops being "sharper" and starts being "airbrushed": mode-seeking with a heavy hand flattens texture into plastic. k is a fine-grained knob (step 0.005) for a reason.
Copying the paper's tsr_sigma 3 onto SDXL. That number belongs to a flow-matching schedule; on an eps model with a karras-type scheduler the same value shifts where the rescaling bites, often to a place you don't want. Default first, then experiment.
Stacking it with an aggressive correction. TSR and a standard-deviation rescale are pulling in opposite directions - one sharpens toward modes, the other flattens magnitude. If you stack them, you'll conclude both are useless. Use one.
The single-slot trap. Another pack's RescaleCFG, Mahiro or RenormCFG node chained after this one owns ComfyUI's single CFG-function slot and TSR silently does nothing. CFG Plan Readout on the model settles it in one queue - and if you want to see the per-step effect rather than guess at it, CFG Measure: Per-Step Probe writes the numbers to output/cfg_probe/.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | — | |
| scale | FLOAT | -1.0-1–100 | The guidance scale w for this rule. -1 uses the sampler's cfg value. |
| k | FLOAT | 0.9500.5–1 | 1 = off (paper SD3 / FLUX 0.93). |
| tsr_sigma | FLOAT | 1.00.1–10 | The rescaling's own sigma (paper SD3 / FLUX 3). |
| space | COMBO | auto (the method's own) | Where the rule is computed. Linear rules give the same image in any space; nonlinear ones do not. 'auto' uses the space the method was published in (noise for most, denoised for APG and the angle rule, velocity for flow models). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |