SeaCache
Free Speed on FLUX, Wan and Hunyuan by Skipping the Boring Steps
- model
- model
Every DiT you run does the same thing thirty times: push the latent through the whole transformer, get a velocity, step. On most of those steps the latent barely moved, so you paid full price to recompute almost what you already had. SeaCache's pitch is that you can stop paying for those.
It's the 2026 entry in the caching-accelerator family - same shelf as TeaCache, TaylorSeer, HiCache and Spectrum - and it's the good kind of free lunch: no extra model, no LoRA, no training, no VRAM cost. The paper is Chung et al., SeaCache: Spectral-Evolution-Aware Cache (CVPR 2026); this pack is chinoll's independent ComfyUI adaptation for native FLUX, HunyuanVideo, Wan 2.1 and MiniMax H3. It's brand new, and unlike the rest of that family it has almost no tuning lore yet.
How it works
It's a MODEL -> MODEL patch node, same shape as a LoRA loader: it clones the model and wraps its forward pass, so nothing upstream changes.
Each step it builds a cheap "monitor" tensor from the block input (on FLUX, the image stream after norm and modulation), runs it through the SEA filter - a separable Wiener-style filter over the token grid, parameterised off the current flow sigma (a = 1 - sigma, b = sigma) - and diffs it against the previous step's filtered monitor with a relative L1. While that accumulated difference stays under rel_l1_thresh, the blocks are skipped: instead of re-running them, the node adds the residual stored on the last real pass (block output minus block input) to the current latent. One addition instead of twenty-odd blocks.
When the change crosses the threshold the blocks run for real and the residual refreshes; the first cacheable step always runs in full. State is per CFG branch, so cond and uncond don't starve each other, and the cache is wiped at the start of every sampling run.
The difference from Spectrum or HiCache: those forecast the hidden feature from a fitted basis and predict it on skipped steps. SeaCache predicts nothing - it reuses the last residual behind a change detector.
The inputs that matter
rel_l1_thresh(default0.2) is the whole speed/quality dial: higher reuses more steps. The README's sane band is0.10–0.30, and it goes to3.0, but up there you're caching nearly everything and getting paste. At0the node hands your model straight back - handy for a same-seed A/B.model_type-autoreads the loaded diffusion model's class and picks for you, which is right 99% of the time. Explicit choices:flux,hunyuan_video,wan2.1,minimax_h3- picking wrong raises rather than quietly degrading.start_percent/end_percent(defaults0/1) - the slice of the schedule where caching is live. Pullingend_percentto around0.85–0.9is the standard trick to protect the refinement tail, where a stale residual shows up first as soft microdetail.cache_device- leave it oncuda.cpuonly buys VRAM, and it slows the FFT and the cache transfers.h3_metric_stride- MiniMax H3 only. It spatially subsamples the video tokens before the SEA FFT so the FFT buffer stays bounded at high resolution; the output itself is never downsampled. Default4, drop to2if you have VRAM to spare.
One output, model, wired onward into your sampler or guider. Put the node immediately after the diffusion-model loader and after any LoRA or model-patch nodes - anything downstream that clones or replaces the model drops the patch.
Installing it
Search SeaCache in ComfyUI Manager, or clone it:
cd ComfyUI/custom_nodes
git clone https://github.com/chinoll/ComfyUI-SeaCache
Restart ComfyUI. That's the whole install - no requirements.txt, no model downloads, because the node uses only PyTorch and ComfyUI's own model classes.
Where people get burned
Don't stack it. No EasyCache, TeaCache or Spectrum on the same model - two gates both deciding to skip steps compound into quality collapse, with no warning.
It only understands native models. Diffusers-style loaders, wrappers running their own denoising loop, VACE/S2V Wan and FLUX Kontext are deliberately off the fast path. Kontext falls back to the native forward silently (its per-token timestep modulation breaks the assumption); the rest just don't work with it.
ComfyUI updates can break it. The pack reimplements ComfyUI's native forward functions per architecture, and its own source warns to keep it updated as those APIs change. Small packs lag: if sampling starts erroring right after a ComfyUI update, disable this node and re-test before blaming anything else.
Licensing is a live question. Upstream SeaCache ships no license file, and the README says to confirm its status before redistributing. The adaptation is GPL-3.0-only.
Realistic workflow: set rel_l1_thresh to 0.2, run your usual 25–30 steps, and compare against a 0.0 run at the same seed. If the images match, nudge to 0.25 and see how far it goes before the fine texture does.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| model | MODEL | Connect after Load Diffusion Model and any LoRA/model patches. | |
| model_type | COMBO | auto | SeaCache supports the upstream FLUX/HunyuanVideo/Wan2.1 architectures plus native MiniMax H3. |
| rel_l1_thresh | FLOAT | 0.200–3 | Higher values reuse more steps and trade more quality for speed. 0 disables SeaCache. |
| start_percent | FLOAT | 0.000–1 | — |
| end_percent | FLOAT | 1.000–1 | — |
| cache_device | COMBO | cuda | CPU saves VRAM but makes the SEA FFT and cache transfers slower. |
| h3_metric_stride | INT | 41–16 | MiniMax H3 only: spatial stride for the SEA metric. Higher values avoid large FFT buffers at high resolutions. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| model | MODEL | — |