🚀AGSoft MiniMax H3 Cache
Free steps on MiniMax H3 — when the video barely moves, this node skips the work
- model
- model
If you've ever watched a 33B MiniMax H3 render churn through step after step while the frame barely moves, this is the node you've been looking for. AGSoft MiniMax H3 Cache sits between your model and the sampler and quietly skips the expensive parts: when both the video and the audio barely change from one step to the next, it replays the last full step's output instead of running the model again. Think of it as an adaptive, per-step DeepCache for H3 - no training, no distilled checkpoint, just free speed on the steps that would've been near-identical anyway.
How it works
The mechanism is surprisingly clean. The node clones your MODEL, then monkey-patches the diffusion model's _forward with a wrapper. Every full step it takes a cheap snapshot of the inputs - the video stream and the audio stream, downsampled to strided fp32 slices so the math is nearly free - and caches the raw output. On the next step, before running the model at all, it computes a sampled relative delta between the current input and the snapshot: mean(|cur−prev|)/(mean(|prev|)+eps). If both deltas are below their thresholds, it returns the cached output and skips the forward pass. Audio gets a veto: even if the picture is static, a voice or lip movement changing enough will force a full run, because nobody wants garbled dialogue to save a few seconds.
A few rails keep it from wrecking your renders. Caching only happens inside a window of the sigma schedule (by default 10–90% of the run), the first warmup_steps are always full, and max_steps caps how many skips can happen consecutively before a forced full run. New sampling run? It detects sigma rising again or a fresh schedule and resets its state. The one output is a patched model that plugs straight into your KSampler's model input - wire it in and let it do its thing.
The inputs that actually matter
profile is the whole game. Any preset except Custom overrides every manual widget below it, which is the #1 trap here: people tweak sliders, wonder why nothing changes, and it's because the preset is silently winning. The presets are:
- Balanced (default) - v 0.120 / a 0.100 thresholds, window 10–90%, warmup 2, max 1 consecutive skip. The safe starting point.
- Visual Fast - looser thresholds, wider window, up to 3 consecutive skips. Fastest, riskiest.
- Dialogue Safe - much tighter thresholds. Pick this when speech and lip-sync matter.
- Action Safe - tightest of all, narrowest window. For fast-cut action where anything could move next.
- Custom - turns on
video_threshold,audio_threshold,start_percent/end_percent,warmup_steps,max_steps, the two*_metric_stridewidgets anddevice.
Start on Balanced, then look at the console. With verbose on (it is by default) the node logs every step as RUN or SKIP with both deltas and real wall time, and prints a summary with theoretical speedup and seconds saved at the end of each run. That's your tuning instrument - if it says ~1.3x, that's a real lunch break back. max_steps above 1 is where the risk lives; raise it only if you're not seeing artifacts.
Installing it
It ships in the comfyui-AGSoft pack. ComfyUI Manager (search "comfyui-AGSoft") or:
cd ComfyUI/custom_nodes
git clone https://github.com/Art-xmaster/comfyui-AGSoft.git
Then restart ComfyUI. The pack's real dependencies are modest - numpy, opencv-python, translators - nothing H3-specific, because the H3 model itself comes from ComfyUI's native MiniMax-H3 support. What you do need is the model: ~42.5GB of weights under the H3 Community License, which geofences out the US, EU, UK and Korea. If you're in one of those regions you're not licensed to run the weights at all - check that before you blame this node for anything.
Where people get burned
Beyond the preset-override trap: bypassing the node cleanly disables caching (it falls back to the original forward pass, no harm done), and cache only operates inside the window, so the opening and closing steps of a render stay full-cost. Don't expect a memory fix - this is a speed play, and the 33B model will still eat whatever VRAM it eats. And if you're coming from Wan or LTX hoping for a one-click speedup on a different model, sorry: this wrapper is written specifically against H3's dual-stream forward. For its own model, though, it's the difference between watching paint dry and watching paint dry slightly faster - and on a model this heavy, that's real.
Gotchas
- Tweaking widgets under a non-Custom profile does nothing - set profile to Custom first.
- Voice artifacts = lower
audio_threshold(0.080–0.100), or just switch to Dialogue Safe. - No speedup = you're outside the cache window (early/late steps), or the deltas genuinely exceed thresholds every step - check the RUN/SKIP log before assuming it's broken.
- The pack's docs live on a Telegram channel (t.me/prompt_by_art), not in the repo, if you want the author's own notes.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | Input MiniMax H3 model. --- Входная модель MiniMax H3. | |
| profile | COMBO | Balanced | Versioned preset. Any profile except Custom OVERRIDES the manual widgets below: Balanced: v 0.120 / a 0.100, window 0.10-0.90, warmup 2, max_steps 1; Visual Fast: v 0.140 / a 0.120, window 0.06-0.94, warmup 2, max_steps 3; Dialogue Safe: v 0.078 / a 0.065, window 0.12-0.88, warmup 3, max_steps 1; Action Safe: v 0.065 / a 0.052, window 0.15-0.85, warmup 3, max_steps 1; Custom: uses the manual widgets. --- Версионированный пресет. Любой профиль кроме Custom ПЕРЕОПРЕДЕЛЯЕТ ручные виджеты ниже: Balanced: v 0.120 / a 0.100, окно 0.10-0.90, warmup 2, max_steps 1; Visual Fast: v 0.140 / a 0.120, окно 0.06-0.94, warmup 2, max_steps 3; Dialogue Safe: v 0.078 / a 0.065, окно 0.12-0.88, warmup 3, max_steps 1; Action Safe: v 0.065 / a 0.052, окно 0.15-0.85, warmup 3, max_steps 1; Custom: использует ручные виджеты. |
| video_threshold | FLOAT | 0.1200–1 | CUSTOM ONLY: sampled relative delta threshold for VIDEO: mean(|cur-prev|)/(mean(|prev|)+eps). Recommended range: 0.100-0.140. --- ТОЛЬКО Custom: порог относительной дельты ВИДЕО: mean(|cur-prev|)/(mean(|prev|)+eps). Рекомендуемый диапазон: 0.100-0.140. |
| audio_threshold | FLOAT | 0.1000–1 | CUSTOM ONLY: sampled relative delta threshold for AUDIO (audio veto). Recommended range: 0.080-0.120; lower = safer voice. --- ТОЛЬКО Custom: порог относительной дельты АУДИО (вето по аудио). Рекомендуемый диапазон: 0.080-0.120; ниже = безопаснее голос. |
| start_percent | FLOAT | 0.100–1 | CUSTOM ONLY: start of the cache window, fraction of the sigma schedule. --- ТОЛЬКО Custom: начало окна кэширования, доля расписания сигм. |
| end_percent | FLOAT | 0.900–1 | CUSTOM ONLY: end of the cache window, fraction of the sigma schedule. --- ТОЛЬКО Custom: конец окна кэширования, доля расписания сигм. |
| warmup_steps | INT | 20–100 | CUSTOM ONLY: first N steps of every run are always full. --- ТОЛЬКО Custom: первые N шагов каждого запуска всегда полные. |
| max_steps | INT | 10–64 | CUSTOM ONLY: maximum CONSECUTIVE skips before a forced full run. 0 disables reuse. Higher = faster but riskier. --- ТОЛЬКО Custom: максимум пропусков ПОДРЯД перед принудительным полным прогоном. 0 отключает переиспользование. Выше = быстрее, но рискованнее. |
| video_metric_stride | INT | 121–1024 | Strided sampling of the video tensor for the metric (every Nth element). Higher = cheaper metric, slightly noisier. --- Strided-выборка видео-тензора для метрики (каждый N-й элемент). Больше = дешевле метрика, чуть шумнее. |
| audio_metric_stride | INT | 61–1024 | Strided sampling of the audio tensor for the metric (every Nth element). Denser than video to protect voice. --- Strided-выборка аудио-тензора для метрики (каждый N-й элемент). Плотнее видео для защиты голоса. |
| device | COMBO | auto | Where the metric snapshots/deltas are computed. auto = on the tensor's device; cpu avoids GPU sync but copies data; cuda forces GPU. --- Где считаются снимки/дельты метрики. auto = на устройстве тензора; cpu избегает GPU-синхронизации, но копирует данные; cuda принудительно GPU. |
| verbose | BOOLEAN | true | Per-step logging (RUN with real wall time / SKIP with both deltas / final run summary). --- Логирование каждого шага (RUN с реальным временем / SKIP с обеими дельтами / итоговая сводка). |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |