LoRA Gate (Quantized)
Let a LoRA work the first two steps, then get out of the way
- lora
- lora
Some LoRAs are only useful for the first couple of steps. The classic case is the "diversity" LoRA on a step-distilled turbo model: Krea 2 Turbo and Z-Image Turbo both have a low-variance problem - same prompt, different seeds, suspiciously similar images - and the community fix is an SDA-style LoRA that jitters the high-sigma trajectory back to life. It works, right up until you notice it also steamrolls your style LoRAs and softens fine texture, because it's still applying on step eight. The obvious answer is "run it early and stop", which on a floating-point model means a sigma-scheduled LoRA node. On a quantized model the LoRA path itself is the fragile part, so ComfyUI-QuantizationToolkit ships its own gate: LoRA Gate (Quantized).
What the node actually does
It's a pass-through. It takes one QUANTIZATION_TOOLKIT_LORA entry - normally the output of LoRA Stack Entry (Quantized), which already carries the file path and strength - and hands the same entry back with an active_steps count attached. No model in, no model out, nothing loaded or patched at this point. All the work happens inside Apply LoRA Stack (Quantized).
That's the whole trick, and it's a tidy one. The apply node splits its entries: ungated ones go through the ordinary patch path in whatever mode you picked (Standard, Stochastic, Dynamic); gated ones are routed to the runtime delta path regardless of the stack mode. At each model evaluation the runtime wrapper reads the sampler's actual sigma schedule, takes the current sigma off the timestep, and drops the entry the moment sigma <= sigmas[active_steps]. So active_steps is a sigma boundary at step index N, not a percentage of the run, and on an 8-step schedule "2" means the first two steps exactly.
Two behaviours worth internalising. Intermediate solver evaluations use the same sigma cutoff, so a solver that evaluates the model twice per step doesn't get a bonus LoRA pass. And every sampling invocation starts a fresh gate: a hires-fix second pass or a partial-denoise pass gets its own first N steps, rather than inheriting what the first pass already burned.
The inputs that matter
- lora (required) - the entry from
LoRA Stack Entry (Quantized). Path and strength come along for the ride; the entry node's ownstrengthstill applies. - active_steps (INT, default 2) - how many intervals of the actual schedule the entry stays active for.
0disables the entry outright, and the code is literal about it: a zero-gated entry with a filename that doesn't even exist will not touch disc. Setting the entry's strength to 0, bypassing it, or unplugging it does the same thing.
One output, lora, same QUANTIZATION_TOOLKIT_LORA type - into Apply LoRA Stack (Quantized). The gate has to sit inline: LoRA Stack Entry (Quantized) -> LoRA Gate (Quantized) -> Apply LoRA Stack (Quantized). Wire an entry straight into the apply node and it simply isn't gated, which is a wiring mistake that produces "the gate does nothing" reports.
The README's worked example is Krea 2 Turbo SDA: use the Comfy-format LoRA, set active_steps to 2, and sample on the author's 8-step schedule.
Installing it
ComfyUI Manager, search ComfyUI Quantization Toolkit (publisher sparknight), or:
cd ComfyUI/custom_nodes
git clone https://github.com/SparknightLLC/ComfyUI-QuantizationToolkit
Restart afterwards. The toolkit has no pip dependencies of its own - pyproject.toml lists an empty dependencies array - but it does need ComfyUI 0.32.0 or newer, a comfy-kitchen new enough to expose the TensorCoreConvRotW4A4Layout and AsymW4A8Int8Layout layouts, an NVIDIA GPU with usable INT8 throughput, and optionally a working Triton install for the alternate INT8 backend. Two naming traps: the project was published as ComfyUI-INT8-Toolkit, and the registry ID and internal node IDs were deliberately frozen, so the folder on disk may show up as ComfyUI-INT8-Fast-Fork even though you cloned the QuantizationToolkit repo. Leave it alone.
Where people get burned
- The LoRA seems on for the whole run. Check the wire order first. If the gate is bypassed, the entry is ungated and behaves like any other entry.
ValueError: LoRA Gate requires a sampler that supplies sample_sigmas in transformer_options.A custom or exotic sampler node isn't forwarding the schedule. Stock samplers do; use one, or don't gate.LoRA Gate requires ordinary linear LoRA patches; unsupported target: …DoRA and convolutional adapters can't be gated at runtime, so the toolkit raises instead of quietly leaving them active for the full run. That's the right call, but it means your DoRA goes on the ungated path or not at all.- You can't save it. Gated entries are runtime-only;
Save Quantized Model (DynamicVRAM Safe)will not bake them into a checkpoint. If the effect has to live in the file, bake stock LoRAs before quantization instead - that's whatEnable Quantization on MODEL's bake option is for. - A console warning about W4A8. Gated entries always try the runtime path, but W4A8 targets fall back to ComfyUI's standard patch path with a warning, since a runtime-delta implementation for that format doesn't exist yet.
- Anything else quantization-shaped. A stock
Load LoRAon a quantizedMODELis the trap this whole pack exists to avoid; use the toolkit's entry/apply nodes and keep the fast path. And if you're on a card that already holds the model comfortably, remember quantization buys you little -int8exists for the cases where the model genuinely doesn't fit.
Inputs (2)
| Name | Type | Default | Description |
|---|---|---|---|
| lora | QUANTIZATION_TOOLKIT_LORA | — | |
| active_steps | INT | 20–10000 | Active for the first N intervals of the sampler's actual sigma schedule. 0 disables this entry. Each sampling invocation starts a fresh gate; intermediate solver evaluations use the same sigma boundary. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| lora | QUANTIZATION_TOOLKIT_LORA | — |