Nodes/ComfyUI-ArchAi3d-Qwen/ArchAi3D_Qwen_Encoder_V3
ComfyUI Node

ArchAi3D_Qwen_Encoder_V3

The encoder with a preset for image-vs-text balance and a recommended CFG

By amir84ferdos·Created 11 months ago·Updated 5 months ago· 70
ArchAi3D_Qwen_Encoder_V3
  • clip
  • vae
  • image1_vl
  • image2_vl
  • image3_vl
  • image1_latent
  • image2_latent
  • image3_latent
  • conditioning
  • latent
  • formatted_prompt
  • recommended_cfg
prompt
system_prompt
conditioning_balance
conditioning_balance_override
manual_context_strength1.00
manual_user_strength1.00
image1_labelImage 1
image2_labelImage 2
image3_labelImage 3
image1_latent_strength1.00
image2_latent_strength1.00
image3_latent_strength1.00
debug_modefalse

By the time a pack is on its third encoder, it's solving a very specific annoyance: people don't want to think about strength numbers. Encoder V3 turns the two strength sliders from V2 into a dropdown of named balances - Image-Dominant through Text-Dominant - and then goes one better: it outputs a recommended CFG value tuned to whichever balance you picked, so you can wire it straight into the KSampler's cfg. It's the encoder for "I know what I want, I just don't want to tune it."

The panel

The new controls, in order of importance:

  • conditioning_balance - the V3 preset dropdown. Pick a named balance (the tooltip confirms the spectrum runs Image-Dominant → Text-Dominant) and the node sets the context/user strength pairing for you. The tooltip also notes it works great with ConditioningAverage, which suggests the intended workflow is averaging this conditioning with something else when you want even finer control.
  • conditioning_balance_override - an optional STRING input. The tooltip says to connect a "Conditioning Balance" node here to override the dropdown; leave it empty and the dropdown rules.
  • manual_context_strength (0–3) and manual_user_strength (0–3) - only used when the preset is "Custom". The extended 0–3 range exists "for extreme conditioning control," i.e. for people who know exactly why they want 2.7.

Everything else is the V2 platform: system_prompt, image1/2/3_label, image1/2/3_latent_strength, debug_mode, and the optional vae + *_vl/*_latent inputs.

Outputs

Four - one more than the other encoders:

  • conditioning - text + vision embeddings with reference latent metadata.
  • latent - the image1 latent, VAEDecode-compatible.
  • formatted_prompt - the ChatML prompt for debugging.
  • recommended_cfg (FLOAT) - ⭐ the new one. A suggested CFG scale in the 2.5–5.5 range based on your balance preset. Connect this to the KSampler's cfg parameter and the image/text balance you picked on the front panel is carried all the way to sampling.

How it works

Underneath, it's the V2 two-stage interpolation (context blend + user blend) - the presets are just named combinations of those two alphas, and "Custom" exposes them directly. The extra reach (0–3 vs 0–1.5) is extrapolation past the normal range, which is legal in the interpolation scheme and lets you push either stage hard when the edit demands it. The recommended CFG is computed from the same preset so the prompt strength and the guidance strength move together - the pack's attempt to remove the "why does it look washed out" post-sampling discovery.

The honest take

V3 is a workflow-comfort upgrade, not a quality upgrade over V2 - same pipeline, same outputs, nicer front panel. It earns its keep in two situations: you're building a workflow to share with non-tuners, or you want the balance-to-CFG wiring. If you're solo and comfortable with two sliders, V2 is cheaper. If you find yourself re-deriving "image-dominant needs lower CFG" from memory every session, V3 is the one that just hands it to you.

Install

Pack install, once:

cd ComfyUI/custom_nodes
git clone https://github.com/amir84ferdos/ComfyUI-ArchAi3d-Qwen.git
cd ComfyUI-ArchAi3d-Qwen && pip install -r requirements.txt

Or ComfyUI Manager → "ArchAi3d Qwen". Restart, find it under ArchAi3d/Qwen/Encoders. As always, you need the Qwen-Image-Edit checkpoint + Qwen-VL CLIP, quantized on consumer GPUs. And recommended_cfg only helps if you actually wire it - it's a FLOAT output, not magic.

CategoryArchAi3d/Qwen/Encoders

Inputs (21)

NameTypeDefaultDescription
clipCLIPQwen-VL CLIP model for encoding text and vision tokens
promptSTRINGText prompt (vision tokens inserted automatically in ChatML format)
system_promptSTRINGOptional system prompt (wrapped in ChatML <|im_start|>system block)
conditioning_balanceCOMBOV3 PRESET: Choose conditioning balance (Image-Dominant → Text-Dominant). Works great with ConditioningAverage!
conditioning_balance_overrideSTRINGOPTIONAL: Connect ⚖️ Conditioning Balance node here to override the preset above. Leave empty to use the dropdown.
manual_context_strengthFLOAT1.000–3CUSTOM ONLY: Manual context strength (only used when preset = Custom). Extended range: 0.0-3.0 for extreme conditioning control
manual_user_strengthFLOAT1.000–3CUSTOM ONLY: Manual user strength (only used when preset = Custom). Extended range: 0.0-3.0 for extreme conditioning control
image1_labelSTRINGImage 1Custom label for Image 1
image2_labelSTRINGImage 2Custom label for Image 2
image3_labelSTRINGImage 3Custom label for Image 3
image1_latent_strengthFLOAT1.000–2Image1 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger)
image2_latent_strengthFLOAT1.000–2Image2 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger)
image3_latent_strengthFLOAT1.000–2Image3 latent strength (1.0=normal, <1.0=weaker, >1.0=stronger)
debug_modeBOOLEANfalseEnable console logging (shows preset values, strengths, shapes)
vaeoptVAEVAE for encoding reference latents (required if using latent images)
image1_vloptIMAGEImage 1 for vision encoder (RGB only, expects correct size)
image2_vloptIMAGEImage 2 for vision encoder (RGB only, expects correct size)
image3_vloptIMAGEImage 3 for vision encoder (RGB only, expects correct size)
image1_latentoptIMAGEImage 1 for reference latent (RGB only, expects correct size)
image2_latentoptIMAGEImage 2 for reference latent (RGB only, expects correct size)
image3_latentoptIMAGEImage 3 for reference latent (RGB only, expects correct size)

Outputs (4)

NameTypeDescription
conditioningCONDITIONINGText+vision embeddings with reference latents metadata attached
latentLATENTImage1 latent in standard format (for VAEDecode or other latent nodes)
formatted_promptSTRINGFinal ChatML-formatted prompt with vision tokens (for debugging)
recommended_cfgFLOAT⭐ NEW: Recommended CFG scale based on preset (2.5-5.5). Connect to KSampler's cfg parameter for optimal image/text balance!