ComfyUI Node
SAM-Audio Video Separate
Separates the sound associated with a white mask region from a native ComfyUI VIDEO value, using the video's embedded audio track.
SAM-Audio Video Separate
- pipeline
- video
- mask
- target
- residual
◄description►
◄seed0►
◄inference_steps32►
◄chunk_duration10.0►
◄chunk_overlap1.0►
Categoryaudio/SAM-Audio
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| pipeline | SAM_AUDIO_PIPELINE | — | |
| video | VIDEO | A native ComfyUI VIDEO containing frames and an embedded audio track. | |
| mask | MASK | White selects the visible object whose sound should be isolated. Use one mask for all frames or one per frame. | |
| description | STRING | Optional text guidance to combine with the visual prompt. | |
| seed | INT | 00–18446744073709550000 | Controls SAM-Audio's initial noise for reproducible separation. |
| inference_steps | INT | 322–128 | Number of midpoint function evaluations. Higher values are slower and may improve quality. |
| chunk_duration | FLOAT | 10.00–3600 | Seconds processed per pass. Use 0 to process the entire clip at once. |
| chunk_overlap | FLOAT | 1.00–60 | Seconds shared by adjacent chunks for a smooth crossfade. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| target | AUDIO | — |
| residual | AUDIO | — |