ComfyUI Node
SAM-Audio Text Separate
Separates a described sound from audio. Short lowercase noun or verb phrases such as 'man speaking' or 'dog barking' best match SAM-Audio training.
SAM-Audio Text Separate
- pipeline
- audio
- target
- residual
◄descriptionman speaking►
◄predict_spansfalse►
◄seed0►
◄inference_steps32►
◄chunk_duration10.0►
◄chunk_overlap1.0►
Categoryaudio/SAM-Audio
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| pipeline | SAM_AUDIO_PIPELINE | — | |
| audio | AUDIO | — | |
| description | STRING | man speaking | A concise lowercase description of the sound to isolate. |
| predict_spans | BOOLEAN | false | Ask supported models to locate non-ambient sound events before separation. Uses additional memory and time. |
| seed | INT | 00–18446744073709550000 | Controls SAM-Audio's initial noise for reproducible separation. |
| inference_steps | INT | 322–128 | Number of midpoint function evaluations. Higher values are slower and may improve quality. |
| chunk_duration | FLOAT | 10.00–3600 | Seconds processed per pass. Use 0 to process the entire clip at once. |
| chunk_overlap | FLOAT | 1.00–60 | Seconds shared by adjacent chunks for a smooth crossfade. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| target | AUDIO | — |
| residual | AUDIO | — |