ComfyUI-DiffAid-Patches
A ComfyUI custom node pack implementing Diff-Aid-inspired inference-time text-conditioning patches for Flux and SDXL models.
Nodes (4)
Patch five Flux blocks, leave the other 52 alone — the prompt-following nudge
A text-adherence dial for ComfyUI's MiniMax H3 (experimental, so start low)
Teaching an SDXL U-Net to actually listen
WAN video that finally lets the text get a word in
ComfyUI-DiffAid-Patches
ComfyUI custom nodes that apply Diff-Aid-inspired inference-time text-conditioning patches to supported diffusion models.
This repository is a practical patch pack for ComfyUI inference, not a paper-exact reproduction of the original Diff-Aid training method. It currently provides four nodes:
- Flux-family Diff-Aid Sparse Patch — for Flux-family MMDiT models exposed through ComfyUI as
double_blocks/single_blocks - WAN Diff-Aid Sparse Patch — for native ComfyUI WAN-family video models exposed as a single
blockslist, including WAN 2.1 and WAN 2.2 layouts - MiniMax H3 Diff-Aid Sparse Patch — for native ComfyUI MiniMax H3 models with packed text, visual, audio, and video rows
- SDXL Diff-Aid Cross-Attention Patch — for SDXL-style cross-attention U-Nets, as an architectural adaptation of the same high-level idea
Paper credit / reference
This repository is based on ideas from:
Binglei Li, Mengping Yang, Zhiyu Tan, Junping Zhang, Hao Li Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation arXiv:2602.13585, 2026
https://arxiv.org/abs/2602.13585
The paper presents a lightweight Aid module for rectified text-to-image diffusion transformers that adaptively modulates textual features per token, per block, and per denoising timestep. It is trained with the backbone model frozen, and the learned modulation coefficients are used at inference time as a plug-in enhancement.
Please cite and credit the paper if you use this repository as part of experiments, implementations, ports, or derivative work.
Suggested citation
@article{li2026diffaid,
title={Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation},
author={Li, Binglei and Yang, Mengping and Tan, Zhiyu and Zhang, Junping and Li, Hao},
journal={arXiv preprint arXiv:2602.13585},
year={2026}
}
What the paper does
The paper’s central claim is that prompt adherence and image quality can be improved by strengthening text-image interaction selectively, rather than with a single global guidance scale.
Its Aid module learns a coefficient tensor α that depends on:
- the current transformer block
- the current denoising timestep
- the current text token features
and uses that to modulate the text features before they participate in attention:
c̃ = c + c ⊙ α
The paper also adds:
- bounded modulation (
tanh) - a gating path to encourage sparsity
- regularization on
α - optional DPO / reward-based optimization during training
The result is a learned plug-in that is lightweight in parameter count but still adaptive and interpretable.
What the paper reports
The paper evaluates Diff-Aid on FLUX and SD 3.5 and reports consistent gains in prompt-following and image-quality metrics.
Main reported results from the paper
On the paper’s reported benchmarks:
- SD 3.5 improves from 9.31 → 9.48 on HPSv3 and from 0.72 → 0.77 on GenEval.
- FLUX improves from 10.42 → 10.71 on HPSv3 and from 0.68 → 0.70 on GenEval.
- The paper also reports gains on HPSv2, ImageReward, and Aesthetic Score for both baselines.
Sparse-enhancement finding relevant to this repo
Besides the full learned method, the appendix reports a simpler FLUX sparse enhancement result by only boosting a small subset of blocks:
- selected FLUX blocks:
{1, 15, 36, 41, 48} - with
α = 0.5on those blocks and0.0elsewhere
For that sparse variant, the paper reports the following FLUX results:
- baseline FLUX: HPSv2 28.53 / HPSv3 10.42 / ImgRwd 0.89 / Aes 6.66
- sparse enhancement: 28.61 / 10.57 / 0.98 / 6.71
- full method: 28.80 / 10.84 / 0.95 / 6.76
That appendix result is the main reason this repository starts with a Flux sparse patch instead of trying to fake the entire trained Aid pipeline.
What this repository implements
This repository does not ship the paper’s trained Aid weights or the Aid MLP itself.
Instead, it implements a reviewable inference-time approximation inspired by the paper:
- select where to apply text-conditioning enhancement
- apply a modulation of the form:
c' = c + c * α
- construct
αfrom user-controlled terms:- a base
strength - a normalized sigma-level window
- optional token-position weighting
- optional conditional-branch filtering
- a base
Important caveat
This is not the same as the paper’s learned adaptive Aid module.
In this codebase:
αis not learnedαis not inferred from the current text features by a trained Aid network- block specificity comes primarily from which blocks are patched, not from a separately learned block-wise coefficient function
So the correct description is:
- paper: trained adaptive block/timestep/token modulation
- this repo: practical inference-time modulation inspired by that idea, with a closer approximation for FLUX sparse enhancement and experimental architecture-specific ports for MiniMax H3, WAN, and SDXL
Why there are multiple nodes
The paper works on rectified text-to-image diffusion transformers and evaluates FLUX and SD 3.5.
Those architectures are not interchangeable in ComfyUI integration terms.
Flux-family node
Flux-family models expose transformer block structure that can be patched directly through ComfyUI’s DiT replacement hooks. That makes a sparse block-selection implementation reasonable.
WAN node
Native ComfyUI WAN-family models do not expose Flux-style double_blocks / single_blocks. They expose a single transformer blocks list and use the same patches_replace["dit"][("double_block", i)] replacement convention inside the WAN block loop.
The WAN node uses that native block-replacement path and modulates the WAN context tensor passed to each selected block. For WAN image/video conditioning paths where ComfyUI exposes clip_fea, the node records the image-context prefix length at runtime and preserves that prefix by default, so only the remaining context tokens are modulated.
This is not paper-validated. The Diff-Aid paper focuses on text-to-image models and explicitly frames text-to-video extension as future work, so the WAN node should be treated as a practical experimental port of the same text-conditioning idea, not as a reproduction of reported paper results.
MiniMax H3 node
Native ComfyUI MiniMax H3 uses one two-dimensional packed hidden tensor shaped [packed_sequence_rows, hidden_size]. Depending on the workflow, that sequence can contain Qwen presentation rows, vision pads, visual conditions, reference image/video rows, reference audio rows, target audio rows, and target video rows.
The native model passes mod_segments entries shaped (start_row, stop_row, modulation_index) into each transformer block. The modality class is modulation_index % 3: class 1 is genuine language, class 0 is visual, and class 2 is audio. The H3 node treats this metadata as authoritative and modulates only class-1 ranges. It preserves vision pads even when they interrupt the initial text presentation span, along with every visual condition, reference, audio, and target-video row.
This H3 adaptation is experimental. It has no trained Aid module and no H3-specific validation in the Diff-Aid paper. Its block selection and strength require empirical testing.
SDXL node
SDXL is a cross-attention U-Net, not a Flux-style MMDiT with double_blocks and single_blocks. So the correct hook point is the UNet cross-attention path rather than Flux block replacement.
That means the SDXL node is not paper-validated. It is an architectural adaptation of the same broad principle: strengthen text-conditioning at specific inference locations.
Repository structure
ComfyUI-DiffAid-Patches/
├── __init__.py
├── nodes.py
├── README.md
├── tests/
├── requirements.txt
└── LICENSE
Installation
Clone into ComfyUI/custom_nodes:
cd ComfyUI/custom_nodes
git clone https://github.com/xmarre/ComfyUI-DiffAid-Patches ComfyUI-DiffAid-Patches
No additional Python packages are required beyond a normal ComfyUI installation.
Restart ComfyUI after installation.
Node 1: Flux-family Diff-Aid Sparse Patch
Type: MODEL -> MODEL
This node targets Flux-family models that expose the expected diffusion transformer structure.
What it patches
The node uses ComfyUI model patch hooks to modify selected Flux transformer blocks:
- double-stream blocks: modulates the
txttensor directly - single-stream blocks: modulates only the text-prefix region inside the merged stream, when the required slice metadata is available
Image latent tokens and reference latent tokens are left untouched. If reference latents are present at runtime, the node logs that it detected them and confirms the text-only modulation scope.
The node wraps the model call to derive a normalized sigma level for the current sampling run. The wrapper composes by wrapping the final model_function, so existing ComfyUI or third-party wrappers still get first chance to adjust call parameters.
Inputs
model— inputMODELenabled— bypass switchblock_presetpaper_sparse_flux_double_only_safepaper_sparse_flux_fullcustom_combined_indices
block_indices— comma-separated 1-based combined block indices, used forcustom_combined_indicesstrength— modulation magnitudesigma_start,sigma_end— normalized sigma-level active window, where1.0is the first/high-noise model call in the current run and0.0is the low-noise endsigma_ramp— soft edge width for the normalized sigma windowtoken_weight_modenonelinearexponential
token_tail— final-token weight for non-nonetoken weightingapply_single_stream— whether selected single-stream blocks should also be patchedcond_only— only modulate conditional rows when ComfyUI exposescond_or_uncond; defaultTrue
Outputs
- patched
MODEL - summary
STRING
Paper sparse preset behavior
The paper’s sparse preset is authored for a canonical FLUX.1 block layout:
- 19 double blocks
- 38 single blocks
- paper combined indices:
1, 15, 36, 41, 48
This repository can remap those 1-based combined indices to the currently loaded Flux-family model.
For a canonical 57-block layout, the full paper indices resolve to:
- double blocks:
0, 14 - single blocks:
16, 21, 28
For Flux.2-like reduced layouts, the remapping is a heuristic. The safer default preset is therefore:
paper_sparse_flux_double_only_safe
That preset remaps and activates only the paper-derived double-stream subset. It avoids implying that the full FLUX.1 sparse block set is known-valid for every Flux-family layout.
Use this only when deliberately testing the full remapped sparse set:
paper_sparse_flux_full
Even with the full preset, single-stream blocks are still only patched when apply_single_stream = True.
Legacy workflows that still contain paper_sparse_flux are accepted as an alias for the full remap path, but new workflows should use the explicit preset names.
Recommended starting settings
Safest starting point for Flux-family / Flux.2 experiments:
block_preset = paper_sparse_flux_double_only_safestrength = 0.5sigma_start = 0.0sigma_end = 1.0sigma_ramp = 0.0token_weight_mode = nonecond_only = Trueapply_single_stream = False
For an early/high-sigma-only taper, use a window such as:
sigma_start = 0.55sigma_end = 1.0sigma_ramp = 0.10
Example chain
Flux model -> ModelSamplingFlux (optional) -> Flux-family Diff-Aid Sparse Patch -> sampler
Place this node after model-loader/model-sampling patches and before the sampler. Avoid stacking several nodes that install model_function_wrapper unless that combination has been tested.
Node 2: WAN Diff-Aid Sparse Patch
Type: MODEL -> MODEL
This node targets native ComfyUI WAN-family video models that expose their transformer stack as a single blocks list. It is intended for WAN 2.1 and WAN 2.2 native ComfyUI model objects.
What it patches
The node installs ComfyUI DiT replacement patches on selected WAN blocks using the ("double_block", index) replacement key used by the native WAN block loop. At each selected block it modulates the context tensor before the block receives it:
context' = context + context * α
For text-to-video paths this normally means the text-conditioning context. For image-to-video / TI2V paths with an image-conditioning prefix, preserve_image_context_prefix = True keeps the detected image prefix untouched and applies the modulation only to the remaining context tokens. If ComfyUI or another wrapper does not expose a positive prefix length at runtime, the node does not guess one and treats the full context as modulated text context.
Inputs
model— inputMODELenabled— bypass switchblock_presetpaper_sparse_flux_remappedcustom_block_indices
block_indices— comma-separated 1-based WAN block indices, used forcustom_block_indicesstrength— modulation magnitudesigma_start,sigma_end— normalized sigma-level active windowsigma_ramp— soft edge width for the normalized sigma windowtoken_weight_modenonelinearexponential
token_tail— final-token weight for non-nonetoken weightingpreserve_image_context_prefix— keep detected WAN image-conditioning context tokens unmodified; defaultTruecond_only— only modulate conditional rows when ComfyUI exposescond_or_uncond; defaultTrue
Outputs
- patched
MODEL - summary
STRING
Preset behavior
The WAN node has no paper-authored WAN block list. The paper_sparse_flux_remapped preset remaps the paper appendix FLUX sparse block list from the canonical 57-block FLUX layout onto the detected WAN blocks length. This gives a deterministic sparse starting point for WAN 2.1 / 2.2 experiments without hard-coding one block count.
For example, if the loaded WAN model has 30 blocks, the FLUX sparse list 1,15,36,41,48 remaps to:
1,8,19,22,25
If the loaded WAN model has 40 blocks, it remaps to:
1,11,25,29,34
Use custom_block_indices when you want exact WAN block indices instead of this remap.
Recommended starting settings
Conservative WAN starting point:
block_preset = paper_sparse_flux_remappedstrength = 0.35sigma_start = 0.0sigma_end = 1.0sigma_ramp = 0.0token_weight_mode = nonepreserve_image_context_prefix = Truecond_only = True
For an early/high-sigma-only test, use a window such as:
sigma_start = 0.55sigma_end = 1.0sigma_ramp = 0.10
Example chain
WAN model -> WAN Diff-Aid Sparse Patch -> sampler
Place this node after the WAN model loader and before the sampler. If you also use another WAN patcher that installs block replacement hooks, order can matter. This node preserves an already-installed replacement patch for the same block by calling it after applying the context modulation.
Node 3: MiniMax H3 Diff-Aid Sparse Patch
Type: MODEL -> MODEL, STRING
This node targets the native ComfyUI MiniMax H3 implementation. It installs replacements on selected entries in the model's single transformer blocks list and applies:
text_rows' = text_rows + text_rows * α
Only ranges identified by native mod_segments metadata as modality class 1 are eligible. Token weighting advances continuously over those genuine language rows, so intervening vision pads do not consume token-weight positions. Missing or malformed metadata is treated as a compatibility error; the node does not guess a text prefix.
Inputs
model— native MiniMax H3MODELenabled— bypass switch; a disabled node returns the original model without cloningblock_indices— comma-separated 1-based H3 transformer block indicesstrength— modulation magnitude in[-1.0, 1.0]sigma_start,sigma_end— normalized sigma-level active windowsigma_ramp— soft shoulder width outside that windowtoken_weight_modenonelinearexponential
token_tail— final linguistic-row weight for non-nonetoken weightingcond_only— modulate only conditional rows when ComfyUI exposescond_or_uncond; defaultTrue
Outputs
- patched
MODEL - diagnostic summary
STRING
Experimental starting settings
block_indices = "1,13,25,37,50"strength = 0.20sigma_start = 0.0sigma_end = 1.0sigma_ramp = 0.0token_weight_mode = nonetoken_tail = 0.35cond_only = True
The current default is an evenly distributed exploratory starting set for the current 50-block native model. It is not derived from the paper's FLUX indices, is not a quality preset, and is validated against the detected model's actual block count at runtime. Benchmark other block groups and strength values before drawing conclusions.
Spectrum-compatible workflow order
Load Diffusion Model
-> MiniMax H3 Diff-Aid Sparse Patch
-> Spectrum Apply MiniMax H3
-> guider / scheduler
The block-replacement chain is preserved when both nodes select the same block, including the final H3 block used by Spectrum for actual-step feature capture. Diff-Aid therefore affects actual transformer evaluations and the history Spectrum fits. Spectrum forecast steps skip transformer execution by design.
For enabled, nonzero H3 patches, Diff-Aid also publishes a small versioned Spectrum compatibility descriptor on the cloned model. The descriptor contains only scalar/configuration metadata and the resolved H3 block indices; it does not retain tensors or model objects. Each model call additionally exposes the normalized sigma value already derived by SharedTimestepWrapper, so Spectrum can use the exact same time coordinate without a second normalization pass or CUDA scalar synchronization.
A full sigma_start=0.0, sigma_end=1.0, sigma_ramp=0.0 window has no interior on/off boundary and therefore requires no compatibility refresh. For a partial hard window with sigma_ramp=0.0, Spectrum can detect the exact inclusive active/inactive transition and promote a would-be forecast on that transition step to one real H3 evaluation. Smooth sigma_ramp>0 shoulders remain continuous and are not turned into extra forced NFEs solely because their gain changes over time.
This interoperability metadata does not make the MiniMax H3 Diff-Aid port paper-validated. It only makes the deterministic runtime modulation explicit to a compatible Spectrum consumer.
H3 validation matrix
Use fixed generation settings and compare all four cases:
| Case | Diff-Aid H3 | Spectrum H3 | |---|---:|---:| | Native baseline | Off | Off | | Diff-Aid H3 only | On | Off | | Spectrum H3 only | Off | On | | Combined | On | On |
Test each relevant generation mode separately:
- T2VA
- image-to-video or keyframe-conditioned generation
- Ref2VA
Hold the seed, prompt, model, resolution or MP setting, duration, sampler, step count, and all conditioning inputs fixed. Check prompt adherence, subject and scene stability, motion, visual artifacts, audio intelligibility, audio artifacts, lip synchronization, general audio-video synchronization, runtime, and peak VRAM.
Node 4: SDXL Diff-Aid Cross-Attention Patch
Type: MODEL -> MODEL
This node targets SDXL-style cross-attention U-Nets.
What it patches
The node installs an attn2 patch and modulates cross-attention conditioning tensors before the attention operation:
context_attn2value_attn2when shape-compatible
The patch can be applied:
- to all input/middle/output stages
- to one specific stage
- or to explicit block targets such as
input:4oroutput:7:1
Inputs
model— inputMODELenabled— bypass switchstage_filterallinputmiddleoutput
block_targets— optional explicit targetsstrengthsigma_start,sigma_endsigma_ramptoken_weight_modenonelinearexponential
token_tailcond_only
Outputs
- patched
MODEL - summary
STRING
Example target strings
input:4, middle:0, output:7
Specific transformer inside a spatial transformer:
output:7:1
Recommended first settings
Start conservatively because this is an out-of-paper port:
stage_filter = allblock_targets = ""strength = 0.20to0.35token_weight_mode = lineartoken_tail = 0.35cond_only = True- full sigma window
Example chain
SDXL model -> SDXL Diff-Aid Cross-Attention Patch -> sampler
How the modulation works in this repo
All nodes use the same runtime modulation family:
α = strength × time_gain × branch_gain × token_gain
c' = c + c * α
where:
time_gaincomes from the normalized sigma-level windowbranch_gainis1for conditional rows and0for unconditional rows whencond_only = Trueand ComfyUI exposes an unambiguouscond_or_uncondlayouttoken_gaincomes from the selected token weighting mode- the resulting
αis broadcast over token embeddings
Normalized sigma window
The normalized sigma level is derived from the model-call timestep/sigma sequence seen by the wrapper during the current sampling run:
1.0means the first/high-noise call in the current run0.0means the low-noise end
This avoids comparing raw ComfyUI timestep/sigma units against normalized UI controls.
Token weighting modes
none— all tokens receive the same scalinglinear— modulation decays linearly from early tokens to later tokensexponential— modulation decays exponentially toward later tokens
This was chosen to preserve the paper’s token-importance intuition in a simple, transparent form, even though it is not the paper’s learned per-token Aid network.
Compatibility
Use the Flux node when
- the model is Flux-family
- the loaded diffusion model exposes
double_blocksandsingle_blocks
Use the WAN node when
- the model is a native ComfyUI WAN-family model
- the located diffusion model exposes
blocks,patch_embedding,head, andforward_orig - the workflow uses WAN 2.1 or WAN 2.2 through ComfyUI’s native WAN path
Use the MiniMax H3 node when
- the model uses ComfyUI's native MiniMax H3 diffusion path
- the detected inner model exposes the H3 packed-sequence projections, token refiner, final layer, sigma shifts, and single
blockslist - text scope must follow native
mod_segmentsclass-1 ranges
Use the SDXL node when
- the model is an SDXL-style cross-attention U-Net
- the diffusion model exposes
input_blocks,middle_block, andoutput_blocks - it is not a Flux-family MMDiT
Do not expect
- the Flux node to work on SDXL or WAN
- the WAN node to work on Kijai/WanVideoWrapper internals unless they expose the same native
blocks/patches_replacepath - the MiniMax H3 node to reproduce trained Diff-Aid or paper-validated H3 results
- the SDXL node to behave like the paper’s SD 3.5 implementation
- any node to reproduce the paper’s trained Aid results exactly
Limitations
-
Not a paper-exact reproduction The paper trains lightweight Aid modules; this repo uses hand-constructed runtime modulation.
-
No trained Aid weights included There is no shipped checkpoint corresponding to the paper.
-
FLUX approximation is closer than MiniMax H3, WAN, or SDXL The Flux sparse patch is directly motivated by the paper’s appendix sparse-enhancement result. The MiniMax H3, WAN, and SDXL nodes are experimental architectural ports.
-
WAN is an experimental text-conditioning port WAN support uses the native WAN block hook where available, but the paper does not evaluate text-to-video models.
-
MiniMax H3 placement is unvalidated The default H3 block list is exploratory. There are no H3 ablations establishing useful blocks, strengths, or sigma windows.
-
Single-stream Flux behavior is more sensitive That is why it is disabled by default in the sparse preset path.
-
Results will be model- and workflow-dependent Especially for non-canonical Flux.2 variants, MiniMax H3 audio-video modes, WAN 2.1 / 2.2 video workflows, and SDXL workflows with additional model patches or LoRAs.
-
The exact minimum supported ComfyUI version has not been pinned This node pack requires a ComfyUI build with model patch replacement hooks, attention patch hooks, and model function wrappers.
Practical guidance
- Start with the double-only Flux sparse preset before trying custom indices.
- Keep single-stream patching off unless you are deliberately testing it.
- Keep
cond_only = Trueunless you intentionally want to modulate unconditional/negative-conditioning rows too. - For MiniMax H3, start with low strength and compare the full validation matrix with fixed generation settings.
- For WAN, start with moderate strength and verify motion/detail stability before increasing it.
- For SDXL, start with low strength and broaden only if the effect is too weak.
- Treat this repo as an experimental inference-time patch pack, not as a claim of reproducing the paper’s published numbers.
Acknowledgements
Credit for the original Diff-Aid method, motivation, analysis, and reported findings belongs to the paper authors:
- Binglei Li
- Mengping Yang
- Zhiyu Tan
- Junping Zhang
- Hao Li
This repository is an independent ComfyUI-oriented implementation inspired by that work.
License
MIT