Extensions/ComfyUI-DiffAid-Patches
ComfyUI Extension

ComfyUI-DiffAid-Patches

A ComfyUI custom node pack implementing Diff-Aid-inspired inference-time text-conditioning patches for Flux and SDXL models.

By xmarre·Created 4 months ago·Updated 7 days ago· 16
xmarre/ComfyUI-DiffAid-Patches
Nodes4
On cloudLocal install
Categorymodel_patches/diffaid
Stars16
Updated7 days ago
Readme

ComfyUI-DiffAid-Patches

ComfyUI custom nodes that apply Diff-Aid-inspired inference-time text-conditioning patches to supported diffusion models.

This repository is a practical patch pack for ComfyUI inference, not a paper-exact reproduction of the original Diff-Aid training method. It currently provides four nodes:

  • Flux-family Diff-Aid Sparse Patch — for Flux-family MMDiT models exposed through ComfyUI as double_blocks / single_blocks
  • WAN Diff-Aid Sparse Patch — for native ComfyUI WAN-family video models exposed as a single blocks list, including WAN 2.1 and WAN 2.2 layouts
  • MiniMax H3 Diff-Aid Sparse Patch — for native ComfyUI MiniMax H3 models with packed text, visual, audio, and video rows
  • SDXL Diff-Aid Cross-Attention Patch — for SDXL-style cross-attention U-Nets, as an architectural adaptation of the same high-level idea

Paper credit / reference

This repository is based on ideas from:

Binglei Li, Mengping Yang, Zhiyu Tan, Junping Zhang, Hao Li Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation arXiv:2602.13585, 2026
https://arxiv.org/abs/2602.13585

The paper presents a lightweight Aid module for rectified text-to-image diffusion transformers that adaptively modulates textual features per token, per block, and per denoising timestep. It is trained with the backbone model frozen, and the learned modulation coefficients are used at inference time as a plug-in enhancement.

Please cite and credit the paper if you use this repository as part of experiments, implementations, ports, or derivative work.

Suggested citation

@article{li2026diffaid,
  title={Diff-Aid: Inference-time Adaptive Interaction Denoising for Rectified Text-to-Image Generation},
  author={Li, Binglei and Yang, Mengping and Tan, Zhiyu and Zhang, Junping and Li, Hao},
  journal={arXiv preprint arXiv:2602.13585},
  year={2026}
}

What the paper does

The paper’s central claim is that prompt adherence and image quality can be improved by strengthening text-image interaction selectively, rather than with a single global guidance scale.

Its Aid module learns a coefficient tensor α that depends on:

  • the current transformer block
  • the current denoising timestep
  • the current text token features

and uses that to modulate the text features before they participate in attention:

c̃ = c + c ⊙ α

The paper also adds:

  • bounded modulation (tanh)
  • a gating path to encourage sparsity
  • regularization on α
  • optional DPO / reward-based optimization during training

The result is a learned plug-in that is lightweight in parameter count but still adaptive and interpretable.


What the paper reports

The paper evaluates Diff-Aid on FLUX and SD 3.5 and reports consistent gains in prompt-following and image-quality metrics.

Main reported results from the paper

On the paper’s reported benchmarks:

  • SD 3.5 improves from 9.31 → 9.48 on HPSv3 and from 0.72 → 0.77 on GenEval.
  • FLUX improves from 10.42 → 10.71 on HPSv3 and from 0.68 → 0.70 on GenEval.
  • The paper also reports gains on HPSv2, ImageReward, and Aesthetic Score for both baselines.

Sparse-enhancement finding relevant to this repo

Besides the full learned method, the appendix reports a simpler FLUX sparse enhancement result by only boosting a small subset of blocks:

  • selected FLUX blocks: {1, 15, 36, 41, 48}
  • with α = 0.5 on those blocks and 0.0 elsewhere

For that sparse variant, the paper reports the following FLUX results:

  • baseline FLUX: HPSv2 28.53 / HPSv3 10.42 / ImgRwd 0.89 / Aes 6.66
  • sparse enhancement: 28.61 / 10.57 / 0.98 / 6.71
  • full method: 28.80 / 10.84 / 0.95 / 6.76

That appendix result is the main reason this repository starts with a Flux sparse patch instead of trying to fake the entire trained Aid pipeline.


What this repository implements

This repository does not ship the paper’s trained Aid weights or the Aid MLP itself.

Instead, it implements a reviewable inference-time approximation inspired by the paper:

  • select where to apply text-conditioning enhancement
  • apply a modulation of the form:
c' = c + c * α
  • construct α from user-controlled terms:
    • a base strength
    • a normalized sigma-level window
    • optional token-position weighting
    • optional conditional-branch filtering

Important caveat

This is not the same as the paper’s learned adaptive Aid module.

In this codebase:

  • α is not learned
  • α is not inferred from the current text features by a trained Aid network
  • block specificity comes primarily from which blocks are patched, not from a separately learned block-wise coefficient function

So the correct description is:

  • paper: trained adaptive block/timestep/token modulation
  • this repo: practical inference-time modulation inspired by that idea, with a closer approximation for FLUX sparse enhancement and experimental architecture-specific ports for MiniMax H3, WAN, and SDXL

Why there are multiple nodes

The paper works on rectified text-to-image diffusion transformers and evaluates FLUX and SD 3.5.

Those architectures are not interchangeable in ComfyUI integration terms.

Flux-family node

Flux-family models expose transformer block structure that can be patched directly through ComfyUI’s DiT replacement hooks. That makes a sparse block-selection implementation reasonable.

WAN node

Native ComfyUI WAN-family models do not expose Flux-style double_blocks / single_blocks. They expose a single transformer blocks list and use the same patches_replace["dit"][("double_block", i)] replacement convention inside the WAN block loop.

The WAN node uses that native block-replacement path and modulates the WAN context tensor passed to each selected block. For WAN image/video conditioning paths where ComfyUI exposes clip_fea, the node records the image-context prefix length at runtime and preserves that prefix by default, so only the remaining context tokens are modulated.

This is not paper-validated. The Diff-Aid paper focuses on text-to-image models and explicitly frames text-to-video extension as future work, so the WAN node should be treated as a practical experimental port of the same text-conditioning idea, not as a reproduction of reported paper results.

MiniMax H3 node

Native ComfyUI MiniMax H3 uses one two-dimensional packed hidden tensor shaped [packed_sequence_rows, hidden_size]. Depending on the workflow, that sequence can contain Qwen presentation rows, vision pads, visual conditions, reference image/video rows, reference audio rows, target audio rows, and target video rows.

The native model passes mod_segments entries shaped (start_row, stop_row, modulation_index) into each transformer block. The modality class is modulation_index % 3: class 1 is genuine language, class 0 is visual, and class 2 is audio. The H3 node treats this metadata as authoritative and modulates only class-1 ranges. It preserves vision pads even when they interrupt the initial text presentation span, along with every visual condition, reference, audio, and target-video row.

This H3 adaptation is experimental. It has no trained Aid module and no H3-specific validation in the Diff-Aid paper. Its block selection and strength require empirical testing.

SDXL node

SDXL is a cross-attention U-Net, not a Flux-style MMDiT with double_blocks and single_blocks. So the correct hook point is the UNet cross-attention path rather than Flux block replacement.

That means the SDXL node is not paper-validated. It is an architectural adaptation of the same broad principle: strengthen text-conditioning at specific inference locations.


Repository structure

ComfyUI-DiffAid-Patches/
├── __init__.py
├── nodes.py
├── README.md
├── tests/
├── requirements.txt
└── LICENSE

Installation

Clone into ComfyUI/custom_nodes:

cd ComfyUI/custom_nodes
git clone https://github.com/xmarre/ComfyUI-DiffAid-Patches ComfyUI-DiffAid-Patches

No additional Python packages are required beyond a normal ComfyUI installation.

Restart ComfyUI after installation.


Node 1: Flux-family Diff-Aid Sparse Patch

Type: MODEL -> MODEL

This node targets Flux-family models that expose the expected diffusion transformer structure.

What it patches

The node uses ComfyUI model patch hooks to modify selected Flux transformer blocks:

  • double-stream blocks: modulates the txt tensor directly
  • single-stream blocks: modulates only the text-prefix region inside the merged stream, when the required slice metadata is available

Image latent tokens and reference latent tokens are left untouched. If reference latents are present at runtime, the node logs that it detected them and confirms the text-only modulation scope.

The node wraps the model call to derive a normalized sigma level for the current sampling run. The wrapper composes by wrapping the final model_function, so existing ComfyUI or third-party wrappers still get first chance to adjust call parameters.

Inputs

  • model — input MODEL
  • enabled — bypass switch
  • block_preset
    • paper_sparse_flux_double_only_safe
    • paper_sparse_flux_full
    • custom_combined_indices
  • block_indices — comma-separated 1-based combined block indices, used for custom_combined_indices
  • strength — modulation magnitude
  • sigma_start, sigma_end — normalized sigma-level active window, where 1.0 is the first/high-noise model call in the current run and 0.0 is the low-noise end
  • sigma_ramp — soft edge width for the normalized sigma window
  • token_weight_mode
    • none
    • linear
    • exponential
  • token_tail — final-token weight for non-none token weighting
  • apply_single_stream — whether selected single-stream blocks should also be patched
  • cond_only — only modulate conditional rows when ComfyUI exposes cond_or_uncond; default True

Outputs

  • patched MODEL
  • summary STRING

Paper sparse preset behavior

The paper’s sparse preset is authored for a canonical FLUX.1 block layout:

  • 19 double blocks
  • 38 single blocks
  • paper combined indices: 1, 15, 36, 41, 48

This repository can remap those 1-based combined indices to the currently loaded Flux-family model.

For a canonical 57-block layout, the full paper indices resolve to:

  • double blocks: 0, 14
  • single blocks: 16, 21, 28

For Flux.2-like reduced layouts, the remapping is a heuristic. The safer default preset is therefore:

paper_sparse_flux_double_only_safe

That preset remaps and activates only the paper-derived double-stream subset. It avoids implying that the full FLUX.1 sparse block set is known-valid for every Flux-family layout.

Use this only when deliberately testing the full remapped sparse set:

paper_sparse_flux_full

Even with the full preset, single-stream blocks are still only patched when apply_single_stream = True.

Legacy workflows that still contain paper_sparse_flux are accepted as an alias for the full remap path, but new workflows should use the explicit preset names.

Recommended starting settings

Safest starting point for Flux-family / Flux.2 experiments:

  • block_preset = paper_sparse_flux_double_only_safe
  • strength = 0.5
  • sigma_start = 0.0
  • sigma_end = 1.0
  • sigma_ramp = 0.0
  • token_weight_mode = none
  • cond_only = True
  • apply_single_stream = False

For an early/high-sigma-only taper, use a window such as:

  • sigma_start = 0.55
  • sigma_end = 1.0
  • sigma_ramp = 0.10

Example chain

Flux model -> ModelSamplingFlux (optional) -> Flux-family Diff-Aid Sparse Patch -> sampler

Place this node after model-loader/model-sampling patches and before the sampler. Avoid stacking several nodes that install model_function_wrapper unless that combination has been tested.


Node 2: WAN Diff-Aid Sparse Patch

Type: MODEL -> MODEL

This node targets native ComfyUI WAN-family video models that expose their transformer stack as a single blocks list. It is intended for WAN 2.1 and WAN 2.2 native ComfyUI model objects.

What it patches

The node installs ComfyUI DiT replacement patches on selected WAN blocks using the ("double_block", index) replacement key used by the native WAN block loop. At each selected block it modulates the context tensor before the block receives it:

context' = context + context * α

For text-to-video paths this normally means the text-conditioning context. For image-to-video / TI2V paths with an image-conditioning prefix, preserve_image_context_prefix = True keeps the detected image prefix untouched and applies the modulation only to the remaining context tokens. If ComfyUI or another wrapper does not expose a positive prefix length at runtime, the node does not guess one and treats the full context as modulated text context.

Inputs

  • model — input MODEL
  • enabled — bypass switch
  • block_preset
    • paper_sparse_flux_remapped
    • custom_block_indices
  • block_indices — comma-separated 1-based WAN block indices, used for custom_block_indices
  • strength — modulation magnitude
  • sigma_start, sigma_end — normalized sigma-level active window
  • sigma_ramp — soft edge width for the normalized sigma window
  • token_weight_mode
    • none
    • linear
    • exponential
  • token_tail — final-token weight for non-none token weighting
  • preserve_image_context_prefix — keep detected WAN image-conditioning context tokens unmodified; default True
  • cond_only — only modulate conditional rows when ComfyUI exposes cond_or_uncond; default True

Outputs

  • patched MODEL
  • summary STRING

Preset behavior

The WAN node has no paper-authored WAN block list. The paper_sparse_flux_remapped preset remaps the paper appendix FLUX sparse block list from the canonical 57-block FLUX layout onto the detected WAN blocks length. This gives a deterministic sparse starting point for WAN 2.1 / 2.2 experiments without hard-coding one block count.

For example, if the loaded WAN model has 30 blocks, the FLUX sparse list 1,15,36,41,48 remaps to:

1,8,19,22,25

If the loaded WAN model has 40 blocks, it remaps to:

1,11,25,29,34

Use custom_block_indices when you want exact WAN block indices instead of this remap.

Recommended starting settings

Conservative WAN starting point:

  • block_preset = paper_sparse_flux_remapped
  • strength = 0.35
  • sigma_start = 0.0
  • sigma_end = 1.0
  • sigma_ramp = 0.0
  • token_weight_mode = none
  • preserve_image_context_prefix = True
  • cond_only = True

For an early/high-sigma-only test, use a window such as:

  • sigma_start = 0.55
  • sigma_end = 1.0
  • sigma_ramp = 0.10

Example chain

WAN model -> WAN Diff-Aid Sparse Patch -> sampler

Place this node after the WAN model loader and before the sampler. If you also use another WAN patcher that installs block replacement hooks, order can matter. This node preserves an already-installed replacement patch for the same block by calling it after applying the context modulation.


Node 3: MiniMax H3 Diff-Aid Sparse Patch

Type: MODEL -> MODEL, STRING

This node targets the native ComfyUI MiniMax H3 implementation. It installs replacements on selected entries in the model's single transformer blocks list and applies:

text_rows' = text_rows + text_rows * α

Only ranges identified by native mod_segments metadata as modality class 1 are eligible. Token weighting advances continuously over those genuine language rows, so intervening vision pads do not consume token-weight positions. Missing or malformed metadata is treated as a compatibility error; the node does not guess a text prefix.

Inputs

  • model — native MiniMax H3 MODEL
  • enabled — bypass switch; a disabled node returns the original model without cloning
  • block_indices — comma-separated 1-based H3 transformer block indices
  • strength — modulation magnitude in [-1.0, 1.0]
  • sigma_start, sigma_end — normalized sigma-level active window
  • sigma_ramp — soft shoulder width outside that window
  • token_weight_mode
    • none
    • linear
    • exponential
  • token_tail — final linguistic-row weight for non-none token weighting
  • cond_only — modulate only conditional rows when ComfyUI exposes cond_or_uncond; default True

Outputs

  • patched MODEL
  • diagnostic summary STRING

Experimental starting settings

  • block_indices = "1,13,25,37,50"
  • strength = 0.20
  • sigma_start = 0.0
  • sigma_end = 1.0
  • sigma_ramp = 0.0
  • token_weight_mode = none
  • token_tail = 0.35
  • cond_only = True

The current default is an evenly distributed exploratory starting set for the current 50-block native model. It is not derived from the paper's FLUX indices, is not a quality preset, and is validated against the detected model's actual block count at runtime. Benchmark other block groups and strength values before drawing conclusions.

Spectrum-compatible workflow order

Load Diffusion Model
-> MiniMax H3 Diff-Aid Sparse Patch
-> Spectrum Apply MiniMax H3
-> guider / scheduler

The block-replacement chain is preserved when both nodes select the same block, including the final H3 block used by Spectrum for actual-step feature capture. Diff-Aid therefore affects actual transformer evaluations and the history Spectrum fits. Spectrum forecast steps skip transformer execution by design.

For enabled, nonzero H3 patches, Diff-Aid also publishes a small versioned Spectrum compatibility descriptor on the cloned model. The descriptor contains only scalar/configuration metadata and the resolved H3 block indices; it does not retain tensors or model objects. Each model call additionally exposes the normalized sigma value already derived by SharedTimestepWrapper, so Spectrum can use the exact same time coordinate without a second normalization pass or CUDA scalar synchronization.

A full sigma_start=0.0, sigma_end=1.0, sigma_ramp=0.0 window has no interior on/off boundary and therefore requires no compatibility refresh. For a partial hard window with sigma_ramp=0.0, Spectrum can detect the exact inclusive active/inactive transition and promote a would-be forecast on that transition step to one real H3 evaluation. Smooth sigma_ramp>0 shoulders remain continuous and are not turned into extra forced NFEs solely because their gain changes over time.

This interoperability metadata does not make the MiniMax H3 Diff-Aid port paper-validated. It only makes the deterministic runtime modulation explicit to a compatible Spectrum consumer.

H3 validation matrix

Use fixed generation settings and compare all four cases:

| Case | Diff-Aid H3 | Spectrum H3 | |---|---:|---:| | Native baseline | Off | Off | | Diff-Aid H3 only | On | Off | | Spectrum H3 only | Off | On | | Combined | On | On |

Test each relevant generation mode separately:

  • T2VA
  • image-to-video or keyframe-conditioned generation
  • Ref2VA

Hold the seed, prompt, model, resolution or MP setting, duration, sampler, step count, and all conditioning inputs fixed. Check prompt adherence, subject and scene stability, motion, visual artifacts, audio intelligibility, audio artifacts, lip synchronization, general audio-video synchronization, runtime, and peak VRAM.


Node 4: SDXL Diff-Aid Cross-Attention Patch

Type: MODEL -> MODEL

This node targets SDXL-style cross-attention U-Nets.

What it patches

The node installs an attn2 patch and modulates cross-attention conditioning tensors before the attention operation:

  • context_attn2
  • value_attn2 when shape-compatible

The patch can be applied:

  • to all input/middle/output stages
  • to one specific stage
  • or to explicit block targets such as input:4 or output:7:1

Inputs

  • model — input MODEL
  • enabled — bypass switch
  • stage_filter
    • all
    • input
    • middle
    • output
  • block_targets — optional explicit targets
  • strength
  • sigma_start, sigma_end
  • sigma_ramp
  • token_weight_mode
    • none
    • linear
    • exponential
  • token_tail
  • cond_only

Outputs

  • patched MODEL
  • summary STRING

Example target strings

input:4, middle:0, output:7

Specific transformer inside a spatial transformer:

output:7:1

Recommended first settings

Start conservatively because this is an out-of-paper port:

  • stage_filter = all
  • block_targets = ""
  • strength = 0.20 to 0.35
  • token_weight_mode = linear
  • token_tail = 0.35
  • cond_only = True
  • full sigma window

Example chain

SDXL model -> SDXL Diff-Aid Cross-Attention Patch -> sampler

How the modulation works in this repo

All nodes use the same runtime modulation family:

α = strength × time_gain × branch_gain × token_gain
c' = c + c * α

where:

  • time_gain comes from the normalized sigma-level window
  • branch_gain is 1 for conditional rows and 0 for unconditional rows when cond_only = True and ComfyUI exposes an unambiguous cond_or_uncond layout
  • token_gain comes from the selected token weighting mode
  • the resulting α is broadcast over token embeddings

Normalized sigma window

The normalized sigma level is derived from the model-call timestep/sigma sequence seen by the wrapper during the current sampling run:

  • 1.0 means the first/high-noise call in the current run
  • 0.0 means the low-noise end

This avoids comparing raw ComfyUI timestep/sigma units against normalized UI controls.

Token weighting modes

  • none — all tokens receive the same scaling
  • linear — modulation decays linearly from early tokens to later tokens
  • exponential — modulation decays exponentially toward later tokens

This was chosen to preserve the paper’s token-importance intuition in a simple, transparent form, even though it is not the paper’s learned per-token Aid network.


Compatibility

Use the Flux node when

  • the model is Flux-family
  • the loaded diffusion model exposes double_blocks and single_blocks

Use the WAN node when

  • the model is a native ComfyUI WAN-family model
  • the located diffusion model exposes blocks, patch_embedding, head, and forward_orig
  • the workflow uses WAN 2.1 or WAN 2.2 through ComfyUI’s native WAN path

Use the MiniMax H3 node when

  • the model uses ComfyUI's native MiniMax H3 diffusion path
  • the detected inner model exposes the H3 packed-sequence projections, token refiner, final layer, sigma shifts, and single blocks list
  • text scope must follow native mod_segments class-1 ranges

Use the SDXL node when

  • the model is an SDXL-style cross-attention U-Net
  • the diffusion model exposes input_blocks, middle_block, and output_blocks
  • it is not a Flux-family MMDiT

Do not expect

  • the Flux node to work on SDXL or WAN
  • the WAN node to work on Kijai/WanVideoWrapper internals unless they expose the same native blocks/patches_replace path
  • the MiniMax H3 node to reproduce trained Diff-Aid or paper-validated H3 results
  • the SDXL node to behave like the paper’s SD 3.5 implementation
  • any node to reproduce the paper’s trained Aid results exactly

Limitations

  1. Not a paper-exact reproduction The paper trains lightweight Aid modules; this repo uses hand-constructed runtime modulation.

  2. No trained Aid weights included There is no shipped checkpoint corresponding to the paper.

  3. FLUX approximation is closer than MiniMax H3, WAN, or SDXL The Flux sparse patch is directly motivated by the paper’s appendix sparse-enhancement result. The MiniMax H3, WAN, and SDXL nodes are experimental architectural ports.

  4. WAN is an experimental text-conditioning port WAN support uses the native WAN block hook where available, but the paper does not evaluate text-to-video models.

  5. MiniMax H3 placement is unvalidated The default H3 block list is exploratory. There are no H3 ablations establishing useful blocks, strengths, or sigma windows.

  6. Single-stream Flux behavior is more sensitive That is why it is disabled by default in the sparse preset path.

  7. Results will be model- and workflow-dependent Especially for non-canonical Flux.2 variants, MiniMax H3 audio-video modes, WAN 2.1 / 2.2 video workflows, and SDXL workflows with additional model patches or LoRAs.

  8. The exact minimum supported ComfyUI version has not been pinned This node pack requires a ComfyUI build with model patch replacement hooks, attention patch hooks, and model function wrappers.


Practical guidance

  • Start with the double-only Flux sparse preset before trying custom indices.
  • Keep single-stream patching off unless you are deliberately testing it.
  • Keep cond_only = True unless you intentionally want to modulate unconditional/negative-conditioning rows too.
  • For MiniMax H3, start with low strength and compare the full validation matrix with fixed generation settings.
  • For WAN, start with moderate strength and verify motion/detail stability before increasing it.
  • For SDXL, start with low strength and broaden only if the effect is too weak.
  • Treat this repo as an experimental inference-time patch pack, not as a claim of reproducing the paper’s published numbers.

Acknowledgements

Credit for the original Diff-Aid method, motivation, analysis, and reported findings belongs to the paper authors:

  • Binglei Li
  • Mengping Yang
  • Zhiyu Tan
  • Junping Zhang
  • Hao Li

This repository is an independent ComfyUI-oriented implementation inspired by that work.


License

MIT