Video Cond Merge
Three conditionings, three very different merge modes
- text_conditioning
- character_conditioning
- hdr_conditioning
- merged_conditioning
- merge_report
Once you have more than one signal to give a video model - the prompt, a character or identity embedding, and an HDR descriptor - you have a merge problem. ComfyUI's conditioning is a list of [tensor, dict] entries, and "combining" that could mean three genuinely different things. Video Cond Merge names all three and lets you pick.
The inputs
text_conditioning is required and is the base: the encoded prompt. Crucially, the output keeps its entries and merges the others into them - this node isn't a merge of equals, it's a merge into a spine.
Then two optional slots: character_conditioning (a character or identity embedding from elsewhere in your graph) and hdr_conditioning (from Video HDR Conditioner). Weights control the second two per mode: text_weight (1), character_weight (0.75), hdr_weight (0.5).
merge_mode is the part people get wrong, so here's what each actually does, in the author's words:
concat (the default) appends the other inputs' tokens after the text tokens, multiplied by their weights. Longer conditioning sequence, model sees everything, attention has to do the reconciling. This is what you want for identity - the character embedding keeps its own tokens rather than being averaged into the prompt.
weighted is a weighted average of the token tensors. Same sequence length, blended meaning. Cheaper, and blunt: a strongly-weighted HDR descriptor will pull on your prompt semantics too, because they've literally been averaged together.
priority keeps the text tokens and only copies missing dict keys from the others. It doesn't touch embeddings at all, which makes it the mode for carrying metadata and flags through rather than fusing meanings. Note the tooltip: in priority, both character_weight and hdr_weight are ignored, and in concat the text tokens stay unscaled too - text_weight only applies to weighted.
Outputs are merged_conditioning and merge_report. Read the report; it tells you what happened, and "nothing happened" is a real possible outcome if the input shapes didn't line up.
When to reach for which
If you're putting a character into a video - which in 2026 usually means some form of reference or identity conditioning rather than a from-scratch LoRA - concat is the honest default. The community's hard-won lesson about multi-subject work transfers directly: identity signals that get averaged with prompt tokens turn into mush, and the fix is keeping them separate and letting attention sort it out. That's concat.
Use weighted when you deliberately want one blended aesthetic - a light HDR push warming the whole look, for example - and accept that it changes the prompt's meaning, not just the metadata.
Use priority when a downstream node needs certain dict keys present and you don't want the embeddings touched at all.
The HDR case is worth a sentence of its own: the conditioning that Video HDR Conditioner produces is partly metadata for the pack's own decode nodes. If all you need is that metadata to travel, merging with huge token weights is wasted work - a modest hdr_weight or priority is friendlier to the model.
Practical notes
concat makes sequences longer, and attention cost scales with length. Three sources concatenated is not free; if a clip suddenly samples slower or drifts in adherence, that's the first suspect.
Watch the weights as multipliers, not as sliders. 0.75 and 0.5 are the ship defaults for good reason - a character embedding at 2.0 will absolutely dominate your prompt.
And check the width question: conditioning from a different text encoder has a different embedding width and cannot be merged. The pack does this check in the T2V pipeline's own character_conditioning input ("skipped if its embedding width differs"), and the same physical limit applies here - mismatched widths are a silent no-op or an error depending on the path, never a clever blend.
Install
- ComfyUI Manager → search Radiance → Install → restart → refresh the browser.
- Or:
cd ComfyUI/custom_nodes
git clone https://github.com/fxtd-studios/radiance.git
cd radiance
python -m pip install -r requirements.txt
Windows portable users: run the pip line with python_embeded\python.exe. No models download for this node; the pack's requirements (OpenEXR, OpenImageIO, OpenColorIO, transformers, diffusers, scipy) are the price of the whole thing, not of this node. If Manager gives you an older build than the README's 3.5.0, run Update or git pull inside custom_nodes/radiance.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| text_conditioning | CONDITIONING | Base conditioning, usually the encoded prompt. The output keeps its entries; the other inputs are merged into them. | |
| merge_mode | COMBO | concat | concat: append the other inputs' tokens (times their weights) after the text tokens. weighted: weighted average of the token tensors. priority: keep the text tokens, only copy missing dict keys from the others. |
| character_conditioningopt | CONDITIONING | Optional conditioning (e.g. a character or identity embedding) merged per merge_mode. | |
| hdr_conditioningopt | CONDITIONING | Optional conditioning (e.g. from RadianceVideoHDRConditioner) merged per merge_mode. | |
| text_weightopt | FLOAT | 1.000–2 | Weight of text_conditioning in weighted mode. concat and priority keep the text tokens unscaled. |
| character_weightopt | FLOAT | 0.750–2 | Multiplier on the character tokens in concat and weighted modes. Ignored in priority mode. |
| hdr_weightopt | FLOAT | 0.500–2 | Multiplier on the HDR tokens in concat and weighted modes. Ignored in priority mode. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| merged_conditioning | CONDITIONING | — |
| merge_report | STRING | — |