ComfyUI Extension: Easycontrol KSampler Compatible
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.
KSampler-compatible Anima EasyControl image conditioning for ComfyUI.
README
ComfyUI-EasyControl-KSamplerCompat
KSampler-compatible Anima EasyControl image conditioning for ComfyUI.
One node — Anima EasyControl (KSampler) — takes MODEL + VAE + IMAGE and
returns a MODEL plus an empty LATENT sized to the conditioning resolution.
The reference-image conditioning rides on the returned model as a set of
ModelPatcher object-patches, so it flows straight into the stock KSampler and
composes with the rest of your graph (schedulers, ControlNet, other adapters).
Wire the LATENT output into the KSampler's latent_image and the generation
matches the (aspect-preserved, ~1MP) cond grid automatically — no manual
EmptyLatentImage sizing, and the two streams share a grid for clean spatial
alignment on tasks like colorization.
This is deliberately unlike the upstream Flux / Qwen EasyControl ComfyUI nodes, which make you load a model through a dedicated loader and sample with a dedicated sampler node. Here there is no custom sampler — just a model patch.
What it's for
EasyControl is a frozen-DiT reference-image adapter for the Anima model. The
flagship use case is colorization: feed a grayscale / manga line page into
the IMAGE socket and a color prompt into your normal text encode, and the DiT
colorizes the reference while respecting its structure.
Install
- Clone into
ComfyUI/custom_nodes/:git clone https://github.com/sorryhyun/ComfyUI-EasyControl-KSamplerCompat - Put your EasyControl checkpoint (
*.safetensorstrained withnetworks.methods.easycontrol, e.g. the colorize adapter) inComfyUI/models/loras/. - Restart ComfyUI.
Wiring
UNETLoader ──► MODEL ─┐ MODEL ──► KSampler ──► ...
├─► Anima EasyControl ────┤ ▲
VAELoader ───► VAE ───┤ (KSampler) └─ LATENT ──┘
LoadImage ──► IMAGE ──┘ ▲
easycontrol_lora = <your ckpt>
strength = 1.0, target_megapixels = 1.0
image— the conditioning image (the grayscale / lineart page for colorization). It is resized to at mosttarget_megapixels(aspect-preserving, VAE/patch-snapped) and VAE-encoded once; the returnedLATENTis sized to that same grid, so the cond stream and the sampled target share a resolution.strength— scales the trained cond effect.1.0= as trained. Raise to push harder toward the reference, lower to loosen.mask(optional) — an inpaint mask for theinpaintadapter. Wire the original image intoimageand a mask here; the node fills the white (selected) region with mid-gray and encodes that as the conditioning image — reproducing what theinpaintadapter trained on (cond = original image with a free-form gray hole). The DiT then regenerates the hole consistent with the surrounding pixels and the prompt. Leave unconnected for adapters that take a pre-built cond image (e.g. colorize).target_megapixels(optional) — max output size in megapixels for the returned latent and the cond encode.1.0(default) keeps generation on the ~1MP distribution Anima was trained at by downscaling anything larger; a source already below the cap is left at its native resolution (never upscaled).0keeps the input's native resolution (still snapped to the VAE/patch grid). Feed theLATENToutput into the KSampler'slatent_image.cond_scale_override(optional) — replace the checkpoint's trainedcond_scaleoutright (beforestrength);0= keep the trained value.
Text prompt: encode it the normal way and feed the KSampler's positive/negative
as usual. For the colorize adapter, the prompt carries color facts the
lineart can't ("pink hair, blue eyes, white dress"); structure comes from the
reference image.
How it works
EasyControl extends each DiT block's self-attention to attend over a
reference-image key/value stream in addition to the target tokens, gated by
a trained per-block scalar bias b_cond. On apply:
- The reference image is VAE-encoded and mapped into the DiT input space
(
process_latent_in). - On the first sampling step, the cond stream is walked once through all blocks
to build a per-block
(cond_k, cond_v)cache (deterministic across steps — the cond t-embedding is fixed att=0). Reused for every step and CFG branch. - Each block's
forwardis replaced (via reversible object-patch) with one that runs extended self-attention[target_k ; cond_k]with theb_condbias on the cond columns. Cross-attention and MLP run baseline.
It targets ComfyUI's native Anima/Cosmos backbone
(comfy/ldm/cosmos/predict2.py) directly — a from-scratch reimplementation of
anima_lora's inference path against ComfyUI's split q_proj/k_proj/v_proj
layout. The trained cond-LoRA delta (a fused D→3D tensor) is sliced into q/k/v
thirds and added onto the split projections; the base weights are numerically
identical between the two DiTs (same pretrained model, fused-vs-split layout),
so this reproduces what training saw. No anima_lora vendoring required — it uses
only ComfyUI's own modules plus the checkpoint tensors.
Checkpoints trained with train_adaln (target-stream AdaLN LoRA — per-block
deltas on the adaln_modulation_{self_attn,cross_attn,mlp} up-projections) are
fully supported: the deltas are merged as ordinary ModelPatcher weight patches
(scaled by strength, composing with lowvram loading, block compile, and other
LoRAs), while the cond-stream prefill subtracts them again so the reference
stream sees the frozen modulation exactly as training did. Channel-scaled
checkpoints (*.inv_scale tensors from anima_lora's per-channel gradient
rebalance) are also applied faithfully. The loader is strict: a checkpoint
carrying tensors this node doesn't implement is refused with an error instead of
silently dropping them (a dropped trained feature degrades output with no
warning — if you hit this, update the node).
Known limitations / notes
- Latent space. The node assumes
VAE.encode→process_latent_inyields the same latent the DiT receives for the noisy target. This is the principled match for the native Anima model; if a future build changes the latent pipeline, the cond stream would need the same change. - Block-compile ordering. If you also use a node that rebuilds the DiT (e.g. Anima block-compile), apply this node after it in the chain. The per-block patches resolve their block from the live diffusion_model each forward (rebuild-tolerant), but as with the other Anima nodes, mixing DiT-rebuilding patches is best avoided.
- Attention. Extended attention uses
scaled_dot_product_attentionwith an additive bias mask (memory-efficient backends handle the bias). This differs from anima_lora's training-time flash-LSE path numerically only at the ulp level. - One reference image. If a batch is fed to
IMAGE, the first image is used.
Relationship to the Anima Adapter Loader
This is a separate repo from
ComfyUI-Anima_lora-Adapter
(LoRA / HydraLoRA / ReFT / FeRA / Soft Tokens). Those add residuals to whole
Linear/block outputs and can ride generic forward hooks; EasyControl reaches
inside self-attention, so it needs this dedicated reimplementation.
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.