ComfyUI Node
Audio Reference Encoder
A ComfyUI node in audio/reference with 12 inputs and 2 outputs.
Audio Reference Encoder
- vae
- audio_a
- audio_b
- audio_c
- audio_d
- combined_latent
- info
◄combination_modeattention►
◄weight_a1.00►
◄weight_b1.00►
◄weight_c1.00►
◄weight_d1.00►
◄extract_segmentsfalse►
◄segment_duration2.00►
Categoryaudio/reference
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| vae | VAE | Stable Audio VAE model | |
| audio_a | AUDIO | First reference audio | |
| combination_mode | COMBO | attention | How to combine multiple references |
| audio_bopt | AUDIO | Second reference audio | |
| audio_copt | AUDIO | Third reference audio | |
| audio_dopt | AUDIO | Fourth reference audio | |
| weight_aopt | FLOAT | 1.000–2 | — |
| weight_bopt | FLOAT | 1.000–2 | — |
| weight_copt | FLOAT | 1.000–2 | — |
| weight_dopt | FLOAT | 1.000–2 | — |
| extract_segmentsopt | BOOLEAN | false | Extract most characteristic segments |
| segment_durationopt | FLOAT | 2.000.5–10 | Duration of segments to extract (seconds) |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| combined_latent | LATENT | — |
| info | LATENT_INFO | — |