DDRK Omega Sampler
Universal domain-adaptive sampler for ComfyUI. Flow Matching + EDM with adaptive phase routing, momentum, SABER stabilization and perceptual sharpening
Nodes (7)
Getting some prompt weight back when `(word:1.4)` stops working
Leave steps and CFG at 0 and it'll tell you what it picked
Your steps and CFG, two dials for everything else
When to reach for DDRK Omega Sampler
The DDRK Omega Scheduler in practice
What model am I even running? DDRK Omega Smart Config knows
One KSampler for Flux, SDXL, and everything in between
DDRK Omega Sampler
Domain-adaptive Diffusion Robust Kernel
One sampler node for Flow Matching and EDM models alike.
</div>Overview
DDRK Omega is a single sampler node that handles Flow Matching models (Flux, SD3, Qwen, Krea, HiDream, Chroma, Lumina, Wan) and EDM models (SDXL, SD 1.5, SD 2) without switching samplers or relearning parameters per family. It asks the model which family it is and recalibrates internally.
The goal is a node you drop in and use — not one you tune for an hour per checkpoint.
Version 1.7.0 added DDRK Omega Lite, a thin preset-driven wrapper for ordinary use. Version 1.9.0 is the first release validated on real images rather than only by CPU tests: it adds a second pass (latent-space hires fix) and fixes Flow Matching noise injection - see Measured on images.
Version 1.10.0 is an audit release. 1.9.0 did not import at all (a stray-line IndentationError); 1.10.0 also fixes EDM img2img / hires fix running on the Flow Matching path, s_churn and momentum_beta numerics, adds an opt-in zero-cost HC2 corrector and a 148-test CPU suite with an exact-solution accuracy bench - see Checked on an exact problem and CHANGELOG.md.
Version 1.11.0 adds DDRK Omega Auto, a one-click node that picks steps, CFG, schedule and integrator for the model it is given, and the HC3 integrator: third order at one model call per step, and self-damping at high CFG, where every other multistep sampler tested overshoots - see How it compares.
Version 1.11.1 is the image A/B of 1.10-1.11: HC3, ddrk_model_beta and the fixed sigma_adapt measured on real generations for the first time, on four checkpoints. Three results changed the presets: ddrk_model_beta is now the Flow Matching schedule of Lite and Auto whenever CFG is above 1.5, sigma_adapt is gone from Lite, and Lite fast on EDM is HC3 - see Measured on images (1.11.1) and the sheets in abresult/.
| | | |:--|:--| | HC3 integrator | Third order at 1 model call/step; falls back towards first order by itself where high CFG makes extrapolation unsafe | | HC2 integrator | Second order at 1 model call/step — Heun quality at roughly half the compute on Flow Matching | | Adaptive phase routing | Three phases, each with its own integrator policy and post-processing | | Ancestral SDE | Calibrated noise split (Karras et al. 2022), gated to flat regions | | Per-step telemetry | Opt-in JSON + CSV + summary logs of every internal decision | | Family auto-detection | Schedules, integrator and enhancer defaults resolved per model family |
</div>Installation
cd ComfyUI/custom_nodes/
git clone https://github.com/HVOSTOVSKY/DDRK-Omega-Sampler.git
Restart ComfyUI. No dependencies beyond what ComfyUI already requires.
Quick start
- Add DDRK Omega Auto (one-click) (right-click → Add Node → sampling → DDRK Omega), or open the ready-made graph from the workflow templates: DDRK Omega Auto - text to image.
- Connect
model,positive,negativeandlatent_imageexactly as you would a KSampler, and VAE Decode after it. - Queue. Leave
stepsandcfgat 0: the node recognises the model and picks them. Thesettingsoutput - connect it to a Preview Any node - says what it chose, e.g.SDXL [EDM] | 25 steps (auto) | CFG 6 (auto) | HC3 integrator, Karras.
Then, only if needed:
- Turbo, Lightning, Hyper, LCM, DMD or any few-step LoRA/finetune: set
model_typeto turbo / lightning / few-step (4-8 steps at CFG 1). Flux schnell and Z-Image Turbo are recognised on their own; a turbo LoRA cannot be seen from the model. - Quality:
fast~0.6x the steps,balancedthe model's usual count,best= 1.5x the steps on SD/SDXL or a second pass at 1.33x resolution on Flow Matching (~1.7x the time). - Your model card says otherwise: type its steps and CFG in; any value above 0 is used as is.
- img2img: lower
denoise.
Use the model's native resolution (about 1 MP for Flux, Krea 2, Qwen, SDXL). Nothing in a sampler recovers what a 512x512 canvas cannot hold - see the first comparison below.
Which node?
| Node | For | You set | |:--|:--|:--| | DDRK Omega Auto (one-click) | Everyone. Plug in and queue | nothing; optionally quality, turbo mode, steps/CFG overrides | | DDRK Omega Lite | Your own steps and CFG, two validated dials | steps, CFG, quality, character | | DDRK Omega Unified KSampler (advanced) | Every control: schedules, integrators, SDE, churn, restarts, enhancers, second pass, telemetry | everything | | DDRK Omega Sampler / Scheduler | SamplerCustom / SamplerCustomAdvanced graphs | SAMPLER / SIGMAS objects |
All nodes live under sampling → DDRK Omega.
Measured on images (1.11.1)
Bench: every variant with the same prompt, seed and resolution, one model per run, each image compared by eye and by RMSE to a reference of the same seed, as in 1.9.0. The reference on Flow Matching is HC2 at 4x the steps (100 model calls for Anima, 48 for Krea 2), on EDM RK4 at the same steps (4x the calls). Seeds 5050 and 505050, the same as in 1.9.0. ComfyUI's own samplers (DPM++ 2M, DPM++ 2M Karras) ran on the same schedule through the same graph. RTX 2070 8 GB. The sheets and every number are in abresult/ (results.csv); the bench is the ablation harness of the Elysium project.
Checkpoints: MolKeunMix Anima (Flow Matching, 25 steps, CFG 4, 832x1216), RedCraft Krea 2 (Flow Matching, distilled, 12 steps, CFG 1, 1024x1024), One Obsession (Illustrious, EDM, 20 steps, CFG 5, 832x1216) and Jib Mix Qwen Image (14 steps, CFG 1, one seed, no reference).
| Claim before 1.11.1 | On images | Verdict |
|:--|:--|:--|
| ddrk_model_beta: 1.7-4.3x closer on FM (exact problem) | Anima, CFG 4: 0.075 / 0.091 against 0.114 / 0.202 for the model's plain schedule. Krea 2, CFG 1: 0.187 / 0.103 against 0.205 / 0.085 | Confirmed at CFG > 1, both seeds; a tie at CFG 1, as the bench predicted. Now the Lite/Auto FM schedule at CFG >= 1.5 |
| sigma_adapt = 0.10 (the fixed 1.8.0 controller) | Anima: 0.193 / 0.304 against 0.114 / 0.202 without it; washed out on one seed, the dress changes colour on the other | Negative. Removed from Lite |
| HC3 beats Heun on EDM at equal calls | Illustrious, 20 calls: HC3 0.099 / 0.167, Euler 0.140 / 0.172, Heun 10 steps 0.167 / 0.190, DPM++ 2M Karras 0.122 / 0.166 | Confirmed. Lite fast on EDM is HC3 now. Heun at 40 calls (0.099 / 0.158) stays balanced |
| HC3 is the most accurate on FM at CFG 4 | Anima: HC3 0.114 / 0.202, HC2 0.118 / 0.183, DPM++ 2M 0.116 / 0.194. Krea 2: HC3 0.205 / 0.085, HC2 0.203 / 0.071 | Not confirmed on images: a tie with HC2 and DPM++ 2M |
| FM at CFG 1: all good samplers tie | Krea 2: Euler, HC2, HC3, HC3 + beta within 0.02 of each other on each seed | Confirmed |
| HC2 shows no advantage on EDM | HC2 0.122 / 0.165 against DPM++ 2M Karras 0.122 / 0.166 | Confirmed |
| sharpness = 0.12 | RMSE unchanged (0.119 / 0.206 against 0.114 / 0.202); edge contrast of HC3 0.92x / 0.74x of the reference goes to 1.14x / 0.93x | A look, not an accuracy dial: it gives back the edges the final jump takes away |
| Second pass adds detail | Anima: detail (mean gradient) 1.05x of the reference against 0.96x; Krea 2: 0.91x against 0.83x; +65% time | Confirmed, with a slightly narrower luminance range on Krea 2 |
ddrk_model_beta at CFG 4: closer to the reference on both seeds. The brooch, the bow at the waist and the hands follow the reference; with the plain schedule the larger final one-shot jump to sigma 0 averages such details into something else.
sigma_adapt = 0.10 moves the image furthest from the reference. It was in Lite balanced and best since 1.7.0 on the strength of a measurement of the pre-1.8.0, sign-inverted controller. The fixed controller, measured on images for the first time, lost on both seeds.
HC3 on EDM at equal model calls. At 20 calls HC3 matched Heun at 40 on seed 5050 and was the closest of the one-call methods on both seeds.
<img src="abresult/04_illustrious_equal_calls_s5050.jpg" width="100%" alt="EDM integrators at equal calls">On Flow Matching HC3, HC2 and DPM++ 2M are a draw on images - the integrator matters less than the schedule there.
<img src="abresult/01_anima_integrators_s5050.jpg" width="100%" alt="FM integrators at equal calls">At CFG 1 on a distilled model the variants tie - they move the composition, not the detail.
<img src="abresult/05_krea_cfg1_s5050.jpg" width="100%" alt="Krea 2 at CFG 1">Not measured here: hc2_free_corrector on its own (it is part of HC3), hc2_space = flow, SDE, SABER and the Auto node as a whole. Flux.2 Klein was in the plan and dropped out: its INT8 Qwen3-8B text encoder loads as float32 (~32 GB), so generation took hours - a text-encoder problem, not a sampler one.
How it compares (1.11.0)
No GPU image run backs this section (for images see above): it is the exact-solution bench (python tests/bench_analytic.py, below) against ComfyUI's own samplers, at equal model calls, on each family's usual schedule. Error to the exact solution of the sampling ODE (lower is better):
| EDM, Karras | 6 calls | 10 | 16 | 25 | |:--|--:|--:|--:|--:| | euler, CFG 1 | 0.151 | 0.089 | 0.055 | 0.035 | | ipndm, CFG 1 | 0.028 | 0.011 | 0.0034 | 0.0012 | | uni_pc_bh2, CFG 1 | 0.187 | 0.035 | 0.0042 | 0.0011 | | DDRK hc2, CFG 1 | 0.048 | 0.033 | 0.016 | 0.0056 | | DDRK hc3, CFG 1 | 0.035 | 0.010 | 0.0026 | 0.0012 | | euler, CFG 6 | 0.31 | 0.13 | 0.092 | 0.068 | | ipndm, CFG 6 | 0.80 | 0.27 | 0.11 | 0.012 | | deis, CFG 6 | 0.51 | 0.16 | 0.013 | 0.0052 | | uni_pc_bh2, CFG 6 | 1.38 | 0.55 | 0.084 | 0.0058 | | DDRK hc2, CFG 6 | 0.70 | 0.22 | 0.025 | 0.010 | | DDRK hc3, CFG 6 | 0.37 | 0.11 | 0.029 | 0.0066 |
Where DDRK is strong
- HC3 is the only sampler that is good in both regimes. At CFG 1 it matches the best ODE samplers in ComfyUI (ipndm, uni_pc) from 10 calls up. At CFG 6 with few calls every extrapolating method overshoots - ipndm and uni_pc end up 2.6-4.5x worse than Euler at 6 calls - while HC3 stays within 20% of Euler and beats everything at 10. Same one model call per step as Euler.
- On Flow Matching at CFG 4 HC3 is the most accurate at 10-25 calls (0.030 at 25 against 0.031-0.052 for the rest).
- One node for both families, with the family read from the model, not guessed.
Where it is not
- Flow Matching at CFG 1: all good samplers tie. With the model's shifted schedule the last step is a one-shot jump to sigma 0 from 0.1-0.4, it returns the posterior mean, and that averaging - not the integrator - is most of the error and most of the lost fine detail. The lever there is the schedule:
ddrk_model_beta(the model's own shift, steps denser at both ends) cut the distance to the data distribution 1.7-4.3x at equal steps on the bench. It is offered, not defaulted, until an image A/B confirms it. - 5-6 calls at CFG 1 on EDM: ipndm, with a four-point history, is still more accurate.
- Stochastic sampling:
dpmpp_3m_sdereproduces the FM data distribution about 2x better than any ODE sampler at 25 calls. A stochastic HC3 was built and measured no better, so it did not ship. DDRK's own maskedsde_strengthwas the worst of all on that measure (74 against 25-27 for plain ODE); treat it as a variation dial and leave it at 0 for quality.
Measured on images (1.9.0)
Bench: one process loads the model once and runs every variant with the same seed, prompt and resolution; each image is compared by eye and by RMSE to a high-accuracy reference of the same seed (RK4, 24 steps = 96 model calls; lower RMSE = closer to the exact solution of the same ODE). Models: RedCraft Krea 2 (FM, distilled, 12 steps, CFG 1, 1024x1024, 3 seeds) and MolKeunMix Anima (FM, 25 steps, CFG 4, 832x1216, 2 seeds). Hardware: RTX 2070 8 GB. All sheets are in docs/ab-1.9.0/.
Canvas size matters more than any sampler setting.
<img src="docs/ab-1.9.0/01_canvas_512_vs_1024.jpg" width="100%" alt="512 vs 1024">The pre-1.8.0 Flow Matching defaults washed images out. Cosine schedule + auto_flow_shift + limiter 0.995 + the old sigma_adapt controller: RMSE to the reference 0.286 / 0.281, against 0.108 / 0.155 for the current defaults.
The second pass adds real detail (1.33x, denoise 0.35): on every seed of both models, without changing the composition. Cost on an RTX 2070: Krea 2 12 steps went from 71 s to 125 s; peak VRAM 7.1 GB at 1360x1360.
<img src="docs/ab-1.9.0/03_second_pass_s505050.jpg" width="100%" alt="second pass">HC2 is the most accurate integrator at equal model calls where it matters (Anima, CFG 4, ~25 calls):
| Integrator | Calls | RMSE seed 5050 | RMSE seed 505050 | |:--|--:|--:|--:| | Euler, 25 steps | 25 | 0.123 | 0.210 | | Heun, 13 steps | 25 | 0.123 | 0.230 | | HC2, 25 steps | 25 | 0.108 | 0.155 |
<img src="docs/ab-1.9.0/04_integrators_equal_calls_s505050.jpg" width="100%" alt="integrators">On the distilled Krea 2 at 12 steps and CFG 1 the three integrators were visually close. The integrator matters where the trajectory is hard.
hc2_space = flow is a draw on images. It is the exact FM parameterisation of HC2's correction (below), and on the analytic FM problem it is more accurate from 10 steps up but less accurate at 6-8. On images it was 0.106 vs 0.108 on one seed and 0.165 vs 0.155 on the other. It ships opt-in; ve stays the default.
What is actually verified
This section exists because it is easy for a project like this to accumulate claims. Image-path results below were measured with the built-in telemetry on real generations across five checkpoints (Anima, Krea2, Qwen, Illustrious, PonyXL); convergence orders were measured separately on the stated synthetic ODE.
Measured
Integrator convergence. On the semi-linear test ODE with a known solution, the measured convergence orders are: Euler 1.02, Heun 2.03, RK4 4.02, HC2 order 2 2.01, and HC2 order 3 3.11. At an equal budget of 32 model calls, HC2 order 2 is 13.8x more accurate than Heun, and HC2 order 3 is 11x more accurate than RK4. These are numerical integration results, not an image-quality score.
HC2 vs Heun on Flow Matching, equal compute. HC2 at 20 steps (20 calls) vs Heun at 10 steps (19 calls), three seeds, every enhancer at zero: HC2's final latent held a 16-25% wider dynamic range on all three seeds, mean +21%. The same metric had previously separated Euler from Heun/RK4 and matched blind visual judgement both times.
HC2 vs Heun on Flow Matching, equal step count. Both at 8 steps on a cosine schedule, three seeds: HC2 matched or beat Heun on 2 of 3 seeds (mean +5% range) while using 9 model calls against 15 and finishing roughly 40% faster.
HC2 does not lean on the enhancer stack. Turning sharpening, SDE and SABER off moved HC2's result by +0.2% while Heun's dropped 4.6%. HC2's detail comes from the integration, not from filters applied afterwards.
Integrator ordering on Flow Matching at equal compute. Euler came last on all three seeds, both by eye and by dynamic range, which ran 12-17% narrower than Heun/RK4.
Stability and reproducibility. No NaN or Inf across every logged run. Identical inputs reproduce identical output; residual variation between runs is ~1e-7 and comes from non-deterministic GPU kernels, not from the sampler.
Measured, and negative
HC2 shows no advantage on EDM. Illustrious, three seeds, equal settings: mean range difference +1.2%, and HC2 used 10% more model calls to get it. On EDM, Heun or auto remains the sensible choice.
A plausible explanation: HC2 assumes the denoiser varies smoothly in lambda = -log(sigma). The Karras EDM schedule is already close to uniform in lambda, so it does much of that work already. Flow Matching schedules cannot be uniform in lambda — sigma reaches exactly zero, so lambda diverges — which is where an exponential integrator has the most to offer. This explanation is plausible but not proven, and it cannot be tested directly on Flow Matching for the same structural reason.
The selective corrector did not show a consistent measurable gain. It remains opt-in and off by default. Adaptive step placement did: sigma_adapt=0.10 was the largest single measured gain (+3.3% dynamic range), while the response saturates above 0.20 and higher values can hurt. That measurement is of the pre-1.8.0 controller, which live telemetry showed was sign-inverted (it lengthened the step after a high-error step, raised the last non-zero sigma by 10% and produced near-empty steps). 1.8.0 fixes the controller. Measured on images in 1.11.1, the fixed controller lost on both seeds (Anima, CFG 4: RMSE 0.193 / 0.304 against 0.114 / 0.202 without it), and Lite no longer uses it.
SABER consistently narrows dynamic range. It is a deliberate smoothing control with a real cost and is not a default in Lite. Repeated external reports of a SABER memory leak were investigated and rejected as false positives; no leak workaround is applied.
Small SDE strengths were indistinguishable from off. Values from 0.08 through 0.15 stayed within 0.2% of the deterministic dynamic range in the measured runs. The old sharpness cap was binding, not a plateau: 0.12 still increased the measured effect, so larger values remain reachable but unvalidated.
Not verified
- Cross-enhancer interactions and AB2 extrapolation have not been fully ablated. SABER, SDE, sharpening, and sigma adaptation do have the individual measurements stated here; unlisted combinations remain starting points rather than tuned optima.
- HC2's third-order mode is mathematically verified on synthetic ODEs but rarely accepted on real trajectories at low step counts, and has not been shown to improve images. HC3 (1.11.0) was measured on images in 1.11.1: better than Euler and Heun at equal calls on EDM, a tie with HC2 on Flow Matching.
- Dynamic range is a proxy metric. It tracked visual quality in every comparison where the difference was obvious, and stopped discriminating on subtle ones.
- GPU image quality is not automatically tested; CPU tests verify numerics, invariants, shape handling, and wrapper equivalence.
Checked on an exact problem (1.10.0)
tests/analytic.py draws every latent element from a Gaussian mixture. For that data the denoiser and the exact solution of the sampling ODE are known in closed form, so the error of a sampler can be measured directly instead of judged. It is a check of the numerics, not of image quality. RMSE to the exact solution at equal model calls, smooth mixture (python tests/bench_analytic.py prints all four cases):
| | EDM, Karras, 12 calls | EDM, 30 calls | FM, model schedule, 12 calls | FM, 30 calls |
|:--|--:|--:|--:|--:|
| Euler | 7.4e-2 | 3.0e-2 | 1.07e-1 | 4.5e-2 |
| Heun (15 steps at 30 calls) | - | 1.27e-2 | - | 4.3e-2 |
| DPM++ 2M (ComfyUI) | 2.1e-2 | 2.4e-3 | 6.2e-2 | 1.75e-2 |
| HC2 | 2.6e-2 | 3.6e-3 | 6.2e-2 | 1.74e-2 |
| HC2 + hc2_free_corrector | 1.3e-2 | 8.8e-4 | 6.3e-2 | 1.66e-2 |
What it established:
- Measured orders (log-uniform schedule): Euler 1.0, Heun 2.0, RK4 4.0, HC2 2.0; HC2 with the free corrector ~3.
momentum_betabroke Heun and RK4 (fixed in 1.10.0): at 0.25 RK4 fell from order 4 to 0.9 and Heun from 2 to ~1.1. On realistic EDM schedules that was 1.1-6x more error; on FM with the model's schedule it happened to be 0-17% less, by partly offsetting the final jump. Momentum now applies to Euler steps only.- The free corrector lowers HC2's EDM error 1.1-2x at 8-12 calls and 1.3-4x at 20-30, at no extra calls. On FM it is within a few percent: there the final one-shot jump to sigma 0 dominates the error - the last step is first order in every sampler - so the integrator matters less than where the last non-zero sigma sits.
sigma_adapt = 0.10was 0-22% less accurate than no adaptation in every case;hc2_space = flowwas a draw, as on images.- HC2 and DPM++ 2M are within a few percent of each other, as expected of two second-order exponential multistep methods. HC2's additions are the limiter, the FM parameterisation, the order selection and now the corrector.
The HC2 integrator
What it is
The core is an exponential (semi-linear) multistep method — the same family as DPM-Solver++(2M) and UniPC. That part is published mathematics, not invented here. The slope limiter and the measured-error order selection built on top are specific to this sampler.
Why it beats Heun per model call
The sampling ODE is dx/dsigma = (x - D(x, sigma)) / sigma. It is linear in x, with all the difficulty living in the denoiser D. Euler on it already integrates the linear part exactly — the update reduces to DDIM. Euler's only error is treating D as constant across the step.
Heun and RK4 spend their extra model calls re-approximating the whole right-hand side, including the linear part that needed no approximating, and they evaluate D at intermediate points on a predicted trajectory, so those evaluations carry their own error.
HC2 keeps the exact linear solve and approximates D as linear in lambda = -log(sigma), using the previous step's evaluation — already exact, already paid for, no prediction required:
x_next = e^(-h) * x + (1 - e^(-h)) * D_n + limit( r * [(h - 1) + e^(-h)] )
h = log(sigma_n / sigma_next) r = (D_n - D_prev) / h_prev
The first two terms are Euler/DDIM. The third is the correction, whose coefficient behaves as h^2/2 for small h — second order, at one model call per step.
The slope limiter
r is an extrapolation from past data. Under high CFG the denoiser swings hard between steps and an unlimited extrapolation overshoots — the classic oscillation of high-order schemes near a sharp feature.
Finite-volume CFD solved this decades ago with slope limiters: keep full accuracy where the solution is smooth, drop toward first order where it is not, decided per element. Here the correction is capped elementwise at limiter_kappa x the magnitude of the first-order step. Unlike a clamp on latent values, this never touches the output range, so it costs no contrast.
In practice the limiter engages on 1-3% of elements on a typical step, rising to ~28% on the first step with history. It intervenes, it does not dominate.
Order selection
With hc2_max_order = 3, HC2 fits a quadratic through three past evaluations and compares the third-order term against the second. If each successive term is smaller, the expansion is behaving and the extra order is used; if the third rivals the second, the fit is being driven by noise in D rather than real curvature, and HC2 falls back to second order. Both the ratio and the order chosen are logged.
Third order also pays one extra model call on the first step. A multistep method's first step is first-order, and that single step otherwise caps the entire run at second order — measured: cold start converged at 4x per halving, seeded start at 7.7x toward the theoretical 8x.
The HC3 integrator (1.11.0)
HC3 is HC2 with two additions. Both are free: one model call per step, like Euler.
1. A corrector that costs nothing (the UniC idea from UniPC, Zhao et al. 2023). HC2 has to extrapolate the denoiser across a step from past values. One step later the model has been evaluated at the step's end anyway, so the finished step can be redone by interpolating between known values - linear through the last two, quadratic through three:
x_n <- x_n + alpha_n * [ phi2(h) r_a + (2 phi3(h) - h phi2(h)) dd ] - (the correction HC2 applied)
r_a = (D_n - D_{n-1}) / h dd = (r_a - r_{n-1}) / (h + h_{n-1})
phi2(h) = h - 1 + e^-h phi3(h) = h^2/2 - h + 1 - e^-h
It is applied as a delta, so anything done to the latent in between (SABER, clamps) is kept, and it switches itself off after SDE noise, churn or a restart jump, where the latent did not arrive by that step. Measured order: 2.9-3.0 on both families (HC2: 2.0).
2. Trust damping. Extrapolation is what makes multistep methods accurate on a resolved trajectory and what makes them overshoot on an unresolved one - high CFG at few steps, where the guided denoiser swings between evaluations. The slope history tells the two apart:
rho = |r_n - r_{n-1}| / (|r_n| + |r_{n-1}|) (norms per image in the batch)
theta = 1 - rho r_n <- theta * r_n
On a smooth trajectory consecutive slopes agree, rho is O(h) and HC3 keeps its order; where they disagree, theta goes to 0 and the step becomes DDIM, which is exact for a denoiser that is constant over the step. DPM-Solver and UniPC lower their order by step index ("lower order final"); HC3 decides per step and per image from what the model actually did. The trust per step is in the telemetry (hc3_trust).
Candidates that were measured and not adopted: a third-order predictor (worse at 16-25 calls), no limiter or a looser one (mixed), squared or linear-gain damping (worse at CFG 1), a stochastic HC3 (no better than dpmpp_3m_sde), and a Richardson correction of the final jump (better at 6-8 FM steps, worse above).
Architecture
PHASE 1 (~55% FM / ~65% EDM of steps)
Integrator chosen per step in auto mode, or pinned
FM ancestral SDE injection, gated to flat regions
EDM optional churn (Karras Alg. 2), SABER at very high sigma
PHASE 2 (~35% FM / ~22% EDM)
Integrator per phase policy
FM SABER spatial fusion
EDM no SABER - mid-step blur shifts anatomy
PHASE 3 (remainder)
HC2 in auto mode
FM final perceptual sharpen on the last step
EDM no sharpen
Adaptive integrator order. In auto mode, phase 1 picks per step from a curvature measure: the fractional change of the denoiser output over the previous step. Because it is dimensionless, one pair of thresholds is valid on both families — an earlier version divided by the sigma step, which made the same quantity read ~0.6 on EDM and ~17 on FM.
Ancestral SDE split. A step from sigma to sigma_next decomposes into a shorter deterministic step to sigma_down plus noise of standard deviation sigma_up, calibrated so the combined variance reproduces sigma_next's marginal exactly (Karras et al. 2022). The spatial gating on top — noise only into flat, low-detail regions — is this sampler's own texture-preservation heuristic, not part of that derivation.
Soft clamp. EDM only, bound tied to the noise level (max(4*sigma, 10)) rather than a constant. A constant bound was clipping ~39% of the tensor at high sigma — ordinary early noise, not divergence. Flow Matching gets no per-step clamp at all; both families keep a wide final guard against actual blowups, which widens with the output sigma so a latent handed on still noisy (SplitSigmas) is not clipped.
Family detection. Inside the sampler the family comes from the model (isinstance(model_sampling, CONST), as in ComfyUI's own samplers), not from the schedule. Up to 1.9.0 it was sigmas.max() > 5, which sent EDM img2img below denoise ~0.6 down the Flow Matching path.
Edge detection. The LoG mask flags pixels whose local curvature exceeds the background level, estimated robustly from the median. A percentile cutoff was used before and was tautological — it marked a fixed 20% of every tensor as "edge" regardless of content, including at step 0 where the latent is still pure noise.
Reproducibility. SDE noise and EDM churn draw from one seeded generator. With sde_seed = -1 the seed derives from the global RNG, which ComfyUI seeds from the workflow seed, so runs repeat from the seed widget alone.
Nodes
| Node | Purpose |
|:--|:--|
| DDRK Omega Auto (one-click) | Recognises the model from its ComfyUI config class and picks steps, CFG, schedule and HC3; settings output says what it chose. Start here. |
| DDRK Omega Lite | Thin preset wrapper over the Unified node. Five adjustable controls: seed, steps, CFG, quality, and character. |
| DDRK Omega Unified KSampler | Full all-in-one replacement with every schedule, integrator, enhancer, and diagnostic control. |
| DDRK Omega Sampler | Returns a SAMPLER object for use with SamplerCustom. |
| DDRK Omega Scheduler | Returns a SIGMAS schedule only. Also exposes beta_a / beta_b. |
| DDRK Omega Smart Config | Introspection: detected family, hint text, guidance_embed, recommended settings. Deliberately does not guess steps or CFG. |
DDRK Omega Lite mapping
Lite calls the same Unified node and the same sample_ddrk_omega implementation; it does not contain a second sampler. ddrk_auto selects the family-appropriate schedule, denoise is 1.0, auto-optimization and stochastic features are off, and all unlisted full-node controls use the values shown below. Tests compare every quality/character combination against the equivalent full-node configuration with torch.equal.
Quality
| Model family | Fast | Balanced (default) | Best | |:--|:--|:--|:--| | Flow Matching | HC2 order 2 | HC2 order 2 | HC2 order 2, second pass 1.33x | | EDM | HC3 | Heun | RK4 |
sigma_adapt is 0 everywhere since 1.11.1 (image A/B above). On Flow Matching the schedule is ddrk_model_beta when CFG is 1.5 or more, the model's own schedule below that; on EDM it is Karras.
Euler is deliberately absent from the Flow Matching row. HC2 order 2 costs exactly the same one model call per step as Euler, and on the analytic convergence problem it was more accurate at every step count tested - 4x better at 2 steps, 21x at 3, 17x at 8, 41x at 20. There is no step count at which Euler is the better trade on FM, so no quality level offers it.
Because of that, the FM quality dial is not a speed control - all three levels cost one model call per step, apart from the second pass at best. The speed control is the steps widget. Since 1.11.1 fast and balanced are the same on Flow Matching: the only difference between them was sigma_adapt, which lost on images.
Since 1.9.0, best on FM is balanced plus the second pass (1.33x upscale, denoise 0.35, about 40% of the base steps and at least 4). It replaced HC2 order 3, which live telemetry showed was accepted on one step in nine and never produced a visible gain. best is therefore the one FM quality level that costs more time - about 1.7x in the measurements above.
HC2 is deliberately absent from the EDM row: controlled testing found no HC2 advantage on EDM. HC3 is fast since 1.11.1: at the same one call per step it was the most accurate one-call method on both seeds of the image A/B, where Euler had been. balanced (Heun, 2 calls per step) and best (RK4, 4) are the higher-compute rungs, and in the A/B Heun at 40 calls was as close to the RK4 reference as anything measured.
Character
| Choice | Sharpness | SABER fusion | Meaning | |:--|--:|--:|:--| | Neutral (default) | 0.00 | 0.00 | No stylistic enhancer | | Sharp | 0.12 | 0.00 | Measured sharpness setting; FM only | | Smooth | 0.00 | 0.10 | Deliberate SABER smoothing with its measured dynamic-range cost |
On EDM, sharp is bit-identical to neutral because sharpening is disabled for that family by design. SABER is never a default: it consistently narrows dynamic range and is exposed only through the explicit smooth choice.
Fixed Lite controls: sde_strength=0, s_churn=0, restart_repeats=0, momentum_beta=0, dyn_thresh_percentile=1.0 (off since 1.8.0), limiter_kappa=1.0, hc2_corrector=0, latent_rescale=0, content_aware=true. (hc2_max_order is set by the quality dial, see above.)
Schedulers
| Scheduler | Family | Notes |
|:--|:--|:--|
| ddrk_auto | both | EDM: ddrk_edm_karras. FM: ddrk_model (since 1.8.0; it used to be ddrk_cosine). Recommended. |
| ddrk_model_beta | both | ComfyUI's beta scheduler (0.6, 0.6) on the model's own sigma table: the model's shift, steps denser at both ends, a much smaller final jump (0.114 vs 0.250 at 10 FM steps). Analytic bench: 1.7-4.3x closer to the data distribution on FM at equal steps. Image A/B (1.11.1): 1.5-2.2x closer to the reference on Anima at CFG 4, a tie at CFG 1. Lite and Auto use it on Flow Matching at CFG >= 1.5. Falls back loudly to ddrk_model above 138 steps. |
| ddrk_model | both | The model's own schedule: comfy.samplers.calculate_sigmas(model_sampling, "simple", steps). The shift comes from the model and any ModelSampling* node in the graph; flow_shift / auto_flow_shift are ignored. Falls back loudly to ddrk_flow_linear / ddrk_edm_karras when no model sampling is available. |
| ddrk_cosine | FM | Cosine decay. Densest at sigma ~1: with a shift on top, the first steps are near-empty and the final jump is large (live Krea 2: 0.37-0.40 to zero at 1 MP). Not used by ddrk_auto any more. |
| ddrk_beta | FM | Shaped by beta_a / beta_b; defaults 2.0 / 1.0 give Karras-like shrinking steps |
| ddrk_flow_linear | FM | Linear with shift |
| ddrk_flow_cosmos | FM | Cosmos-style tail |
| ddrk_fewstep | FM | Aggressive shift for 4-8 steps |
| ddrk_edm_karras | EDM | Karras rho=7 |
| ddrk_edm_poly | EDM | Polynomial tail |
| ddrk_edm_simple | EDM | Exponential |
Parameters
Core
| Parameter | Effect |
|:--|:--|
| integrator | hc3 (default for new nodes since 1.11.0) / hc2 / auto / rk4 / heun / euler. At <=6 steps this is forced to euler unless HC2 or HC3 is selected; the console reports any override. |
| sde_strength | Ancestral SDE amount. FM only — silently ignored on EDM, which uses s_churn. |
| sharpness | Final-step perceptual sharpen. FM only — sharpening is disabled on EDM by design. |
| saber_fusion | Spatial/temporal stabilization weight. 0 disables the module entirely. |
| momentum_beta | Adams-Bashforth 2 slope extrapolation for Euler steps only (1.0 = full AB2, second order at one call per step). 0 = plain Euler. Ignored by Heun, RK4 and HC2: on a higher-order method it reduced it to first order, see above. |
| dyn_thresh_percentile | Percentile latent limiter. 1.0 = off, the default since 1.8.0. Engages only below 40% (FM) / 30% (EDM) of sigma_max. At 0.995 it clipped the final FM latent in 4 of 6 live Krea 2 runs (final max = -min exactly); on EDM it fires on most steps. |
| latent_rescale | Attenuates values beyond ~2 std from the per-image mean. Not classical CFG-rescale. 0 = off. |
HC2
| Parameter | Effect |
|:--|:--|
| limiter_kappa | Slope limiter strength. 1.0 means the correction may at most double or cancel the step, never reverse it. Lower to 0.5-0.7 if high CFG still blows out highlights. |
| hc2_max_order | 2 (default) or 3. Third order uses two past evaluations and one extra call to bootstrap. |
| hc2_corrector | Experimental. 0 = off. Otherwise spends a second call on steps whose correction exceeds this fraction of the first-order step. No measurable effect in testing. |
| hc2_free_corrector | Opt-in, 1.10.0. Zero-cost corrector (UniPC's UniC idea): the model output HC2 needs for the next step anyway is also used to redo the finished step by interpolation. No extra calls. Analytic bench: 1.1-4x less error on EDM, within a few percent on FM. Not yet validated on images. |
| hc2_space | ve (default) or flow. flow uses the exact Flow Matching parameterisation - lambda = log((1-sigma)/sigma) and a (1-sigma) weight on the correction. Measured as a draw on images; no effect on EDM. |
| sigma_adapt | 0 = off. Moves intermediate sigmas to equalise estimated error; step count, start and terminal zero unchanged. Since 1.8.0: a step with above-average HC2 activity shortens the next step, the last non-zero sigma is never raised, and no adapted step is shorter than half its reference step. The +3.3% measurement predates this fix; the fixed controller lost on images in 1.11.1 (see "Measured, and negative") and no preset uses it. |
Second pass (1.9.0)
| Parameter | Effect |
|:--|:--|
| refine_scale | 1.0 = off (default). Above 1.0 the finished latent is upscaled by this factor (bislerp, even sizes, 4-D and 5-D latents) and re-sampled with the same settings. 1.25-1.33 is the measured range. |
| refine_denoise | How much of the upscaled latent is rewritten. 0.35 measured; 0.25 is gentler; above ~0.5 it starts to redraw. |
| refine_steps | Model calls spent at the larger size. 4-8 is enough. |
The second pass needs VRAM for the larger canvas. It was measured on 8 GB for Krea 2 and Anima; very large models (Qwen Image 20B) at 1 MP may not fit at 1.33x.
EDM
| Parameter | Effect |
|:--|:--|
| s_churn, s_tmin, s_tmax, s_noise | Karras Alg. 2 churn, gamma = min(s_churn / steps, sqrt(2) - 1) as in k-diffusion. EDM only; s_churn = 0 disables. Up to 1.9.0 gamma divided by sigma instead and sat at its cap on almost every step. |
Other
| Parameter | Effect |
|:--|:--|
| saber_mode | auto treats a 5D latent with F>1 as video. |
| use_ema_saber, ema_decay | Temporal EMA for video SABER. |
| sde_seed | -1 derives the seed from the global RNG, reproducible from the workflow seed. |
| content_aware | Edge-gated SABER fusion. Turn off for pixel art and flat-shaded styles. |
| auto_optimize | Caps or disables enhancers by compute budget on FM, and at <=10 steps turns auto into hc2. Does not touch steps, CFG or the schedule. |
| smart_defaults | Sets scheduler, shift, integrator and enhancer levels from the detected family. Does not touch steps or CFG. |
| debug_mode, debug_tag | Per-step telemetry to disk. Off by default, zero overhead when off. |
A note on two names.
dynamic_thresholdclips peaks but deliberately omits the renormalization step of Imagen-style dynamic thresholding, because dividing the whole latent by the percentile shifts global contrast on every step it fires.latent_rescalecannot be classical CFG-rescale: ComfyUI combines the conditional and unconditional predictions before the sampler is called, so the two tensors that method needs are not available here.
Telemetry
Enable debug_mode and the sampler writes three files to your ComfyUI output folder: a JSON with everything, a CSV with one row per step, and a plain-text summary.
Per step it records the phase, the integrator chosen, curvature, model calls, which of SDE / SABER / sharpen / churn / clamp / thresholding fired and how far each moved the tensor, HC2's order, limiter activity and error ratio, latent statistics, NaN/Inf flags and wall time. The run header records every resolved parameter, the full sigma schedule, and — from the Unified node — cfg, seed, denoise and scheduler.
This is the most useful part of the project for anyone modifying it. Most of what 1.6.0 fixes was found by reading these logs, not by reading the code: a parameter that never fired, a mask reporting the same value on every step, a step count one lower than requested.
Suggested starting points
Starting points, not tuned optima. The simplest start is the Auto node, which applies these per model. Only the integrator guidance comes from controlled comparisons (image A/B for HC2 on FM, the exact-solution bench for HC3).
Flow Matching
| Parameter | Value |
|:--|:--|
| Steps / CFG | Checkpoint-dependent. Distilled: 4-8 / ~1. Base: 20-40 / 1-4. |
| Scheduler | ddrk_model_beta at CFG > 1 (image A/B 1.11.1); ddrk_auto at CFG ~1 |
| Integrator | hc3 or hc2 - a tie on images |
| SDE strength | 0 (it measured as a variation dial, not a quality dial) |
| Sharpness | 0.10-0.15 |
| SABER fusion | 0.00-0.15 |
EDM
| Parameter | Value |
|:--|:--|
| Steps / CFG | 20-30 / 7-8 base; 4-8 / 1-2 turbo |
| Scheduler | ddrk_auto -> ddrk_edm_karras; ddrk_model for Turbo/Lightning |
| Integrator | hc3: at equal model calls (19-39) it was 1.3-21x more accurate than Heun on the exact problem, CFG 1 and 6; on images (1.11.1) the closest one-call method on both seeds |
| s_churn | 0, or 5-15 to try |
| Sharpness | No effect on EDM |
| SABER fusion | 0 (it narrows dynamic range); 0.20 for a deliberately softer look |
Video (5D latents) — saber_mode = video or auto, saber_fusion 0.20-0.35, ema_decay 0.7.
Known limitations
- RK4 costs 4 model calls per step, Heun 2, HC2 and Euler 1. Compare at equal call count, not equal steps.
- <=6 steps forces euler unless HC2 is selected, and EDM <=10 steps turns rk4 into heun - both regardless of
auto_optimize. The console reports the replacement and the telemetry header recordsintegrator_requestednext to the integrator actually used. sharpnessdoes nothing on EDM. The parameter is shared across families; the feature is not.- HC2 shows no advantage on EDM at equal compute.
- Steps and CFG are never auto-detected, except by the Auto node, and there they are starting values per model class (the ones model makers and ComfyUI's templates use), not measurements. A finetune or a LoRA - above all a turbo/lightning one, which cannot be seen from the model - can need other values; the
settingsoutput shows what was chosen, and steps/CFG above 0 override it. - HC3 and
ddrk_model_betaare validated on images on four checkpoints (1.11.1), not on every family. Video (Wan), SD3/SD3.5, HiDream, Chroma and Lumina were not in the A/B; there Auto keeps the model's own schedule only at CFG ~1. The free corrector is validated only as part of HC3. - SDE noise depends on batch shape.
torch.randn_likedraws a whole tensor from one seeded generator stream. Changing batch size or shape changes how that stream is partitioned, so an item sampled alone is not guaranteed the same SDE noise it receives inside a larger/differently shaped batch. This does not affect deterministic runs with SDE/churn/restarts off. - The second pass costs time and VRAM in proportion to
refine_scalesquared; it is not tested on 20B-class models on 8 GB. - CPU tests do not replace GPU image validation. The automated suite covers schedules, convergence, stepping, batch isolation, 4D/5D execution, Lite and Auto mappings, bit-exact wrapper equivalence, and end-to-end runs on tiny real SD 1.5 and Flux models. Output-quality claims still require controlled GPU generations.
Running the tests
The suite runs on CPU against a real ComfyUI checkout. Installed as a custom node (ComfyUI/custom_nodes/DDRK-Omega-Sampler) it finds ComfyUI on its own; anywhere else, point COMFYUI_PATH at one:
pip install pytest
COMFYUI_PATH=/path/to/ComfyUI python -m pytest # 179 tests, ~30 s
COMFYUI_PATH=/path/to/ComfyUI python tests/bench_analytic.py # comparison tables, ~40 s
Node-level tests replace comfy.sample.sample_custom with a stand-in that skips conditioning and model loading but runs ComfyUI's real KSAMPLER, so noise scaling, the inpaint wrapper and inverse noise scaling are the production code. tests/test_real_models.py goes further: SD 1.5 and Flux built by comfy.supported_models at a few hundred thousand random weights, sampled through the real sample_custom - CFGGuider, conditioning, model management and all. GitHub Actions runs the same suite on every push (.github/workflows/tests.yml).
Changelog
v1.11.1
- Image A/B of 1.10-1.11 on Anima, Krea 2, Illustrious and Qwen Image: sheets and numbers in
abresult/, summary in Measured on images (1.11.1). - Changed (output changes): Lite and Auto use
ddrk_model_betaon Flow Matching at CFG >= 1.5 (Anima, CFG 4: RMSE to the reference 0.075 / 0.091 against 0.114 / 0.202). Lite no longer usessigma_adapt(it lost on images on both seeds);fastandbalancedare therefore the same on Flow Matching. Litefaston EDM is HC3 instead of Euler, at the same cost. - Lite at CFG < 1.5, the Unified node and every schedule/integrator path are unchanged.
v1.11.0
- Added: DDRK Omega Auto (one-click) - recognises the model and picks steps, CFG, schedule and integrator; turbo/few-step mode;
settingsoutput; example workflow in the templates browser. - Added: HC3 integrator - HC2 + zero-cost corrector + trust damping: third order at one call per step, robust at high CFG. Default integrator of new Unified/Sampler nodes and of Smart Config on FM.
- Added:
ddrk_model_betaschedule (opt-in). - Fixed: SD3/SD3.5 were detected as Flux; EDM-type models with an unknown
image_model(Cosmos, PixArt, Hunyuan-DiT) as Flow Matching. - Changed: new Unified/Sampler nodes start neutral (no SDE, sharpening, SABER or momentum); all nodes under sampling → DDRK Omega with descriptions; the example workflows moved to
example_workflows/. - Every 1.10.0 path is bit-identical (918 configurations checked). Details: CHANGELOG.md.
v1.10.0
- Fixed: 1.9.0 failed to import (
IndentationError), so no node loaded. - Fixed: EDM runs starting below sigma 5 (img2img, hires fix, second pass, SplitSigmas) took the Flow Matching path; the family now comes from the model.
- Fixed:
s_churnfollows Karras Alg. 2 (s_churn / steps);momentum_betaapplies to Euler only (it made Heun/RK4 first order); the final clamp no longer clips noisy hand-offs; the second pass keepsnoise_mask/batch_index; FM noise honours the model'snoise_scale; the bundled workflow matches the current nodes. - Changed:
autoshares HC2 history across integrators; Smart Config pickshc2on Flow Matching. - Added: opt-in
hc2_free_corrector;hc2_spaceon the Sampler node; 148 CPU tests, an exact-solution accuracy bench and a CI workflow. - The default FM path (Lite, pinned HC2) is bit-identical to 1.9.0. Details and every measurement: CHANGELOG.md.
v1.9.0
- Second pass (
refine_scale/refine_denoise/refine_steps) on the Unified node; Litebeston FM now uses it instead of HC2 order 3. hc2_space: opt-in exact Flow Matching parameterisation of HC2's correction.- Fixed: FM SDE and FM restart jumps used variance-exploding noise formulas; they now land exactly on the FM marginal.
auto_optimizeon FM at <=10 steps picks HC2 instead of Euler forauto.- First release with GPU image A/B results; see Measured on images and CHANGELOG.md.
v1.8.0
- New
ddrk_modelschedule (the model's own shifted "simple" schedule);ddrk_autoon Flow Matching now uses it instead ofddrk_cosine. sigma_adaptcontroller fixed: sign, last-sigma guard, minimum step.dyn_thresh_percentiledefaults to 1.0 (off) in every node and in Lite.- Honest integrator replacement messages and
integrator_requestedin telemetry. - Found from live Elysium telemetry; details in CHANGELOG.md. Sampler output changes for FM
ddrk_auto, for any run that relied on the old limiter default, and forsigma_adapt > 0.
v1.7.0
- Added DDRK Omega Lite, a thin wrapper with quality and character presets mapped to the full node's validated controls.
- Added full-node/Lite bit-exact tests across FM and EDM mappings, and expanded the shipping-backup batch=1 equivalence guard to four 4D/5D cases with the full enhancer stack.
- Removed an unused private SDE generator alias and corrected the
sde_seed=-1tooltip to match its reproducible implementation. - Documented exact convergence measurements, the EDM HC2 negative result, SABER's measured cost, and SDE's batch-shape limitation.
- No sampler output behavior changed.
v1.6.0
New — HC2 integrator. Exponential multistep, second order at one model call per step, with a slope limiter and optional measured-error third order. See The HC2 integrator.
Crashes and dead features
ddrk_betaraisedNotImplementedError— PyTorch's Beta distribution has noicdf. Replaced with the closed-form Kumaraswamy inverse CDF. Default shape changed from (2, 5), whose last step covered 42% of the sigma range, to (2, 1), which shrinks monotonically.beta_a/beta_bare now reachable from the UI; previously no caller passed them.denoise = 0.0was selectable and divided by zero before sampling started.ddrk_animaappeared in both scheduler dropdowns with no implementation behind it. Removed from the UI, still accepted from saved workflows.sharpnessnever applied on any model — the final-step test compared the loop index againsttotal_steps - 1, which the loop never reached.
Schedules and stepping
- Flow Matching schedules ended with a duplicated zero, so a 20-step request ran 19 steps and an 8-step request ran 7. Phase 3's budget was computed including the phantom step, which cost it a step at 20 and eliminated it entirely at 8.
Numerics
- Curvature is now dimensionless. It previously carried units of 1/sigma, averaging 0.60 on EDM and 17.2 on FM while being compared against fixed cutoffs.
rk4no longer divides by the sigma floor on the terminal step; Euler is exact at that endpoint anyway.- EDM soft clamp bound is
max(4*sigma, 10)instead of a constant 10. - The LoG edge mask thresholds against median-estimated background curvature instead of a fixed percentile, which had reported exactly 0.1976 mean coverage on every step of every model.
dynamic_threshold's early-out is now the exact algebraic no-op condition rather than a hard-coded magnitude.- Fixed-gain sigmoid gates replaced with a scale-invariant z-score gate.
Reproducibility
- EDM churn drew from the global RNG while the SDE had its own generator — and with
sde_seed = -1that generator was created unseeded, so SDE runs were not reproducible at all. Both now share one generator, seeded from the global RNG. - The LoG subsample for large latents now uses a fixed-seed generator.
Naming and honesty
cfg_rescale->latent_rescale; the old key is still accepted.dynamic_thresholddocumented as a percentile limiter, not Imagen dynamic thresholding.- Integrator overrides print what they changed instead of silently discarding an explicit choice.
- SABER is skipped, and reported as not fired, when
saber_fusionis 0. - EDM profiles no longer recommend a
sharpnessvalue. - FM profile integrator changed from a hard
eulertoauto. - Enhancer caps key off the estimated model-call budget rather than the step index, so equal-cost configurations get equal treatment.
Performance
- Curvature is computed only when the integrator is
autoand only in phase 1, removing a GPU->CPU sync per step otherwise. - Triplicated integrator dispatch collapsed into one function.
Earlier
| Version | Summary |
|:--|:--|
| v1.5.2 | AB2 extrapolation replacing EMA momentum, ancestral SDE split, z-score gating, s_tmin/s_tmax |
| v1.5.0 | Architecture introspection, smart_defaults, auto_optimize, per-step FM clamp removed |
| v1.4.x | EDM path tuning, churn, multi-scale sharpen, content-aware SABER |
| v1.3 | SABER 5D crash fix for single-frame video latents |
| v1.2 | 5D reshape fix, LRU cache bounds, sde_seed exposed |
| v1.1 | Removed invalid FSAL-RK4, Karras rho=7 schedule, momentum scope fix |
| v1.0 | Original prototype: hybrid RK4/Heun, SABER, SWT sharpening |
Credits
Concept, prototype and direction — HVOSTOVSKY. The phase-based sampler design, SDE noise masking and SABER stabilization are his.
Production hardening (v1.1-v1.5.2) — iterative audits and implementation by Kimi (Moonshot AI).
Telemetry, correctness pass and HC2 (v1.6.0), live-telemetry fixes (v1.8.0), GPU A/B bench and second pass (v1.9.0), audit, test suite and exact-solution bench (v1.10.0), HC3 and the Auto node (v1.11.0), image A/B of 1.10-1.11 (v1.11.1) — by Claude (Anthropic).
HC2's core follows the exponential-integrator line of work — Lu et al., DPM-Solver++ (2022) and Zhao et al., UniPC (2023). The ancestral noise split follows Karras et al., Elucidating the Design Space of Diffusion-Based Generative Models (2022).
Issues
Open an issue with the model name, your settings, and either the traceback or comparison images. For quality problems rather than crashes, enable debug_mode and attach the JSON — it makes most problems diagnosable without guesswork.
<div align="center">
MIT License
</div> *«One sampler to rule them all — from Flux to SDXL.»*