ComfyUI Extension: ComfyUI-DiffAid-Patches
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.
A ComfyUI custom node pack implementing Diff-Aid-inspired inference-time text-conditioning patches for Flux and SDXL models.
Looking for a different extension?
Custom Nodes (0)
README
ComfyUI-DiffAid-Patches
ComfyUI custom nodes that apply Diff-Aid-inspired inference-time text-conditioning patches to supported diffusion models.
This repository is a practical patch pack for ComfyUI inference, not a paper-exact reproduction of the original Diff-Aid training method. It currently provides three nodes:
- Flux-family Diff-Aid Sparse Patch — for Flux-family MMDiT models exposed through ComfyUI as
double_blocks/single_blocks - WAN Diff-Aid Sparse Patch — for native ComfyUI WAN-family video models exposed as a single
blockslist, including WAN 2.1 and WAN 2.2 layouts - SDXL Diff-Aid Cross-Attention Patch — for SDXL-style cross-attention U-Nets, as an architectural adaptation of the same high-level idea
Paper credit / reference
This repository is based on ideas from:
Binglei Li, Mengping Yang, Zhiyu Tan, Junping Zhang, Hao Li Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation arXiv:2602.13585, 2026
https://arxiv.org/abs/2602.13585
The paper presents a lightweight Aid module for rectified text-to-image diffusion transformers that adaptively modulates textual features per token, per block, and per denoising timestep. It is trained with the backbone model frozen, and the learned modulation coefficients are used at inference time as a plug-in enhancement.
Please cite and credit the paper if you use this repository as part of experiments, implementations, ports, or derivative work.
Suggested citation
@article{li2026diffaid,
title={Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation},
author={Li, Binglei and Yang, Mengping and Tan, Zhiyu and Zhang, Junping and Li, Hao},
journal={arXiv preprint arXiv:2602.13585},
year={2026}
}
What the paper does
The paper’s central claim is that prompt adherence and image quality can be improved by strengthening text-image interaction selectively, rather than with a single global guidance scale.
Its Aid module learns a coefficient tensor α that depends on:
- the current transformer block
- the current denoising timestep
- the current text token features
and uses that to modulate the text features before they participate in attention:
c̃ = c + c ⊙ α
The paper also adds:
- bounded modulation (
tanh) - a gating path to encourage sparsity
- regularization on
α - optional DPO / reward-based optimization during training
The result is a learned plug-in that is lightweight in parameter count but still adaptive and interpretable.
What the paper reports
The paper evaluates Diff-Aid on FLUX and SD 3.5 and reports consistent gains in prompt-following and image-quality metrics.
Main reported results from the paper
On the paper’s reported benchmarks:
- SD 3.5 improves from 9.31 → 9.48 on HPSv3 and from 0.72 → 0.77 on GenEval.
- FLUX improves from 10.42 → 10.71 on HPSv3 and from 0.68 → 0.70 on GenEval.
- The paper also reports gains on HPSv2, ImageReward, and Aesthetic Score for both baselines.
Sparse-enhancement finding relevant to this repo
Besides the full learned method, the appendix reports a simpler FLUX sparse enhancement result by only boosting a small subset of blocks:
- selected FLUX blocks:
{1, 15, 36, 41, 48} - with
α = 0.5on those blocks and0.0elsewhere
For that sparse variant, the paper reports the following FLUX results:
- baseline FLUX: HPSv2 28.53 / HPSv3 10.42 / ImgRwd 0.89 / Aes 6.66
- sparse enhancement: 28.61 / 10.57 / 0.98 / 6.71
- full method: 28.80 / 10.84 / 0.95 / 6.76
That appendix result is the main reason this repository starts with a Flux sparse patch instead of trying to fake the entire trained Aid pipeline.
What this repository implements
This repository does not ship the paper’s trained Aid weights or the Aid MLP itself.
Instead, it implements a reviewable inference-time approximation inspired by the paper:
- select where to apply text-conditioning enhancement
- apply a modulation of the form:
c' = c + c * α
- construct
αfrom user-controlled terms:- a base
strength - a normalized sigma-level window
- optional token-position weighting
- optional conditional-branch filtering
- a base
Important caveat
This is not the same as the paper’s learned adaptive Aid module.
In this codebase:
αis not learnedαis not inferred from the current text features by a trained Aid network- block specificity comes primarily from which blocks are patched, not from a separately learned block-wise coefficient function
So the correct description is:
- paper: trained adaptive block/timestep/token modulation
- this repo: practical inference-time modulation patch inspired by that idea, with a closer approximation for FLUX sparse enhancement and a best-effort SDXL port
Why there are multiple nodes
The paper works on rectified text-to-image diffusion transformers and evaluates FLUX and SD 3.5.
Those architectures are not interchangeable in ComfyUI integration terms.
Flux-family node
Flux-family models expose transformer block structure that can be patched directly through ComfyUI’s DiT replacement hooks. That makes a sparse block-selection implementation reasonable.
WAN node
Native ComfyUI WAN-family models do not expose Flux-style double_blocks / single_blocks. They expose a single transformer blocks list and use the same patches_replace["dit"][("double_block", i)] replacement convention inside the WAN block loop.
The WAN node uses that native block-replacement path and modulates the WAN context tensor passed to each selected block. For WAN image/video conditioning paths where ComfyUI exposes clip_fea, the node records the image-context prefix length at runtime and preserves that prefix by default, so only the remaining context tokens are modulated.
This is not paper-validated. The Diff-Aid paper focuses on text-to-image models and explicitly frames text-to-video extension as future work, so the WAN node should be treated as a practical experimental port of the same text-conditioning idea, not as a reproduction of reported paper results.
SDXL node
SDXL is a cross-attention U-Net, not a Flux-style MMDiT with double_blocks and single_blocks. So the correct hook point is the UNet cross-attention path rather than Flux block replacement.
That means the SDXL node is not paper-validated. It is an architectural adaptation of the same broad principle: strengthen text-conditioning at specific inference locations.
Repository structure
ComfyUI-DiffAid-Patches/
├── __init__.py
├── nodes.py
├── README.md
├── requirements.txt
└── LICENSE
Installation
Clone into ComfyUI/custom_nodes:
cd ComfyUI/custom_nodes
git clone https://github.com/xmarre/ComfyUI-DiffAid-Patches ComfyUI-DiffAid-Patches
No additional Python packages are required beyond a normal ComfyUI installation.
Restart ComfyUI after installation.
Node 1: Flux-family Diff-Aid Sparse Patch
Type: MODEL -> MODEL
This node targets Flux-family models that expose the expected diffusion transformer structure.
What it patches
The node uses ComfyUI model patch hooks to modify selected Flux transformer blocks:
- double-stream blocks: modulates the
txttensor directly - single-stream blocks: modulates only the text-prefix region inside the merged stream, when the required slice metadata is available
Image latent tokens and reference latent tokens are left untouched. If reference latents are present at runtime, the node logs that it detected them and confirms the text-only modulation scope.
The node wraps the model call to derive a normalized sigma level for the current sampling run. The wrapper composes by wrapping the final model_function, so existing ComfyUI or third-party wrappers still get first chance to adjust call parameters.
Inputs
model— inputMODELenabled— bypass switchblock_presetpaper_sparse_flux_double_only_safepaper_sparse_flux_fullcustom_combined_indices
block_indices— comma-separated 1-based combined block indices, used forcustom_combined_indicesstrength— modulation magnitudesigma_start,sigma_end— normalized sigma-level active window, where1.0is the first/high-noise model call in the current run and0.0is the low-noise endsigma_ramp— soft edge width for the normalized sigma windowtoken_weight_modenonelinearexponential
token_tail— final-token weight for non-nonetoken weightingapply_single_stream— whether selected single-stream blocks should also be patchedcond_only— only modulate conditional rows when ComfyUI exposescond_or_uncond; defaultTrue
Outputs
- patched
MODEL - summary
STRING
Paper sparse preset behavior
The paper’s sparse preset is authored for a canonical FLUX.1 block layout:
- 19 double blocks
- 38 single blocks
- paper combined indices:
1, 15, 36, 41, 48
This repository can remap those 1-based combined indices to the currently loaded Flux-family model.
For a canonical 57-block layout, the full paper indices resolve to:
- double blocks:
0, 14 - single blocks:
16, 21, 28
For Flux.2-like reduced layouts, the remapping is a heuristic. The safer default preset is therefore:
paper_sparse_flux_double_only_safe
That preset remaps and activates only the paper-derived double-stream subset. It avoids implying that the full FLUX.1 sparse block set is known-valid for every Flux-family layout.
Use this only when deliberately testing the full remapped sparse set:
paper_sparse_flux_full
Even with the full preset, single-stream blocks are still only patched when apply_single_stream = True.
Legacy workflows that still contain paper_sparse_flux are accepted as an alias for the full remap path, but new workflows should use the explicit preset names.
Recommended starting settings
Safest starting point for Flux-family / Flux.2 experiments:
block_preset = paper_sparse_flux_double_only_safestrength = 0.5sigma_start = 0.0sigma_end = 1.0sigma_ramp = 0.0token_weight_mode = nonecond_only = Trueapply_single_stream = False
For an early/high-sigma-only taper, use a window such as:
sigma_start = 0.55sigma_end = 1.0sigma_ramp = 0.10
Example chain
Flux model -> ModelSamplingFlux (optional) -> Flux-family Diff-Aid Sparse Patch -> sampler
Place this node after model-loader/model-sampling patches and before the sampler. Avoid stacking several nodes that install model_function_wrapper unless that combination has been tested.
Node 2: WAN Diff-Aid Sparse Patch
Type: MODEL -> MODEL
This node targets native ComfyUI WAN-family video models that expose their transformer stack as a single blocks list. It is intended for WAN 2.1 and WAN 2.2 native ComfyUI model objects.
What it patches
The node installs ComfyUI DiT replacement patches on selected WAN blocks using the ("double_block", index) replacement key used by the native WAN block loop. At each selected block it modulates the context tensor before the block receives it:
context' = context + context * α
For text-to-video paths this normally means the text-conditioning context. For image-to-video / TI2V paths with an image-conditioning prefix, preserve_image_context_prefix = True keeps the detected image prefix untouched and applies the modulation only to the remaining context tokens. If ComfyUI or another wrapper does not expose a positive prefix length at runtime, the node does not guess one and treats the full context as modulated text context.
Inputs
model— inputMODELenabled— bypass switchblock_presetpaper_sparse_flux_remappedcustom_block_indices
block_indices— comma-separated 1-based WAN block indices, used forcustom_block_indicesstrength— modulation magnitudesigma_start,sigma_end— normalized sigma-level active windowsigma_ramp— soft edge width for the normalized sigma windowtoken_weight_modenonelinearexponential
token_tail— final-token weight for non-nonetoken weightingpreserve_image_context_prefix— keep detected WAN image-conditioning context tokens unmodified; defaultTruecond_only— only modulate conditional rows when ComfyUI exposescond_or_uncond; defaultTrue
Outputs
- patched
MODEL - summary
STRING
Preset behavior
The WAN node has no paper-authored WAN block list. The paper_sparse_flux_remapped preset remaps the paper appendix FLUX sparse block list from the canonical 57-block FLUX layout onto the detected WAN blocks length. This gives a deterministic sparse starting point for WAN 2.1 / 2.2 experiments without hard-coding one block count.
For example, if the loaded WAN model has 30 blocks, the FLUX sparse list 1,15,36,41,48 remaps to:
1,8,19,22,25
If the loaded WAN model has 40 blocks, it remaps to:
1,11,25,29,34
Use custom_block_indices when you want exact WAN block indices instead of this remap.
Recommended starting settings
Conservative WAN starting point:
block_preset = paper_sparse_flux_remappedstrength = 0.35sigma_start = 0.0sigma_end = 1.0sigma_ramp = 0.0token_weight_mode = nonepreserve_image_context_prefix = Truecond_only = True
For an early/high-sigma-only test, use a window such as:
sigma_start = 0.55sigma_end = 1.0sigma_ramp = 0.10
Example chain
WAN model -> WAN Diff-Aid Sparse Patch -> sampler
Place this node after the WAN model loader and before the sampler. If you also use another WAN patcher that installs block replacement hooks, order can matter. This node preserves an already-installed replacement patch for the same block by calling it after applying the context modulation.
Node 3: SDXL Diff-Aid Cross-Attention Patch
Type: MODEL -> MODEL
This node targets SDXL-style cross-attention U-Nets.
What it patches
The node installs an attn2 patch and modulates cross-attention conditioning tensors before the attention operation:
context_attn2value_attn2when shape-compatible
The patch can be applied:
- to all input/middle/output stages
- to one specific stage
- or to explicit block targets such as
input:4oroutput:7:1
Inputs
model— inputMODELenabled— bypass switchstage_filterallinputmiddleoutput
block_targets— optional explicit targetsstrengthsigma_start,sigma_endsigma_ramptoken_weight_modenonelinearexponential
token_tailcond_only
Outputs
- patched
MODEL - summary
STRING
Example target strings
input:4, middle:0, output:7
Specific transformer inside a spatial transformer:
output:7:1
Recommended first settings
Start conservatively because this is an out-of-paper port:
stage_filter = allblock_targets = ""strength = 0.20to0.35token_weight_mode = lineartoken_tail = 0.35cond_only = True- full sigma window
Example chain
SDXL model -> SDXL Diff-Aid Cross-Attention Patch -> sampler
How the modulation works in this repo
All nodes use the same runtime modulation family:
α = strength × time_gain × branch_gain × token_gain
c' = c + c * α
where:
time_gaincomes from the normalized sigma-level windowbranch_gainis1for conditional rows and0for unconditional rows whencond_only = Trueand ComfyUI exposes an unambiguouscond_or_uncondlayouttoken_gaincomes from the selected token weighting mode- the resulting
αis broadcast over token embeddings
Normalized sigma window
The normalized sigma level is derived from the model-call timestep/sigma sequence seen by the wrapper during the current sampling run:
1.0means the first/high-noise call in the current run0.0means the low-noise end
This avoids comparing raw ComfyUI timestep/sigma units against normalized UI controls.
Token weighting modes
none— all tokens receive the same scalinglinear— modulation decays linearly from early tokens to later tokensexponential— modulation decays exponentially toward later tokens
This was chosen to preserve the paper’s token-importance intuition in a simple, transparent form, even though it is not the paper’s learned per-token Aid network.
Compatibility
Use the Flux node when
- the model is Flux-family
- the loaded diffusion model exposes
double_blocksandsingle_blocks
Use the WAN node when
- the model is a native ComfyUI WAN-family model
- the located diffusion model exposes
blocks,patch_embedding,head, andforward_orig - the workflow uses WAN 2.1 or WAN 2.2 through ComfyUI’s native WAN path
Use the SDXL node when
- the model is an SDXL-style cross-attention U-Net
- the diffusion model exposes
input_blocks,middle_block, andoutput_blocks - it is not a Flux-family MMDiT
Do not expect
- the Flux node to work on SDXL or WAN
- the WAN node to work on Kijai/WanVideoWrapper internals unless they expose the same native
blocks/patches_replacepath - the SDXL node to behave like the paper’s SD 3.5 implementation
- any node to reproduce the paper’s trained Aid results exactly
Limitations
-
Not a paper-exact reproduction The paper trains lightweight Aid modules; this repo uses hand-constructed runtime modulation.
-
No trained Aid weights included There is no shipped checkpoint corresponding to the paper.
-
FLUX approximation is closer than WAN or SDXL The Flux sparse patch is directly motivated by the paper’s appendix sparse-enhancement result. The WAN and SDXL nodes are best-effort architectural ports.
-
WAN is an experimental text-conditioning port WAN support uses the native WAN block hook where available, but the paper does not evaluate text-to-video models.
-
Single-stream Flux behavior is more sensitive That is why it is disabled by default in the sparse preset path.
-
Results will be model- and workflow-dependent Especially for non-canonical Flux.2 variants, WAN 2.1 / 2.2 video workflows, and SDXL workflows with additional model patches or LoRAs.
-
The exact minimum supported ComfyUI version has not been pinned This node pack requires a ComfyUI build with model patch replacement hooks, attention patch hooks, and model function wrappers.
Practical guidance
- Start with the double-only Flux sparse preset before trying custom indices.
- Keep single-stream patching off unless you are deliberately testing it.
- Keep
cond_only = Trueunless you intentionally want to modulate unconditional/negative-conditioning rows too. - For WAN, start with moderate strength and verify motion/detail stability before increasing it.
- For SDXL, start with low strength and broaden only if the effect is too weak.
- Treat this repo as an experimental inference-time patch pack, not as a claim of reproducing the paper’s published numbers.
Acknowledgements
Credit for the original Diff-Aid method, motivation, analysis, and reported findings belongs to the paper authors:
- Binglei Li
- Mengping Yang
- Zhiyu Tan
- Junping Zhang
- Hao Li
This repository is an independent ComfyUI-oriented implementation inspired by that work.
License
MIT
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.