BytePlus Seed Audio 1.0
Two minutes of dialogue, music and foley from one prompt
- AUDIO
- subtitles_json
- srt
- duration
- url
Most TTS nodes read text. This one stages a scene. You describe the voices, the emotion, the pacing, the room, the background music and the sound effects - and then you write the lines, naming characters inline - and you get back up to two minutes of produced audio. Dialogue with two speakers. A voiceover with a music bed and a door slam in the right place. It's the one node in the pack where the prompt is doing sound design rather than synthesis.
The pack's Sound Design and Podcast Clip templates both lean on it, and that's the honest framing: it's for scenes, not for reading a paragraph of documentation aloud. If you just need narration, use Seed Speech TTS - it's more controllable, line by line.
How it works
text_prompt carries everything, up to 3000 characters, and it's worth writing like a script rather than a sentence. Name your characters, put their lines in the prompt, and describe the environment around them. Write it in the same language as the lines you want spoken.
reference_mode is the structural choice, and it changes what else appears on the node:
text only- everything comes from the prompt. No references, no@Audiotags.audio reference- up to three connected clips (reference_audio_1…_3, each up to 30 s), or a speaker ID / cloned voice ID / audio URL in the matchingref_audio_N_sourcefields. You refer to them in the prompt as@Audio1,@Audio2,@Audio3. Slots must be filled from 1 upward; gaps are errors.image reference- one character image (reference_image, orref_image_urlas a link), and the model derives a voice from it. Can't be combined with reference audio.preset voice- pick a built-in TTS 2.0 voice frompreset_voiceand it reads the prompt. No clone, no tags.
Then the mix controls: sample_rate, speech_rate, loudness_rate (−50 to 100, 100 = 2x), pitch_rate (±12 semitones), and seed - which, as everywhere in this pack, only decides whether the node re-runs. model defaults to seed-audio-1.0 and covers 20 languages (English, Chinese, Japanese, Korean, both Spanishes, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish).
The timestamp syntax is the feature that makes this useful for video: a quoted line can start with a range, like [5.5s:8.0s] Wait for me!, and that controls when it's spoken and how long it takes. Synchronising a line to a shot becomes arithmetic instead of trial and error.
Outputs are AUDIO, subtitles_json, srt, duration and url - and every one of them is a list, because of generation_count. Set enable_subtitle if you want the timestamps; audio_format picks what's requested from the API (WAV, MP3, OGG Opus, PCM), though it's decoded to AUDIO either way.
One field group worth knowing about even if you never touch it: aigc_watermark adds the audible AI marker, and aigc_metadata plus content_producer / produce_id / content_propagator / propagate_id write provenance into the audio header. If you're producing something that has to declare itself, that's where.
Install and the key
cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
Restart (ComfyUI 0.31.0+), or install through Manager by searching BytePlus ModelArk. Speech is a separate product with a separate key: create it in the Seed Speech console (activate the services first) and save it as the Seed Speech key in Settings → BytePlus, or BYTEPLUS_SEED_SPEECH_API_KEY in user/.env. Seed Speech is ap-southeast-1 only.
Where people get burned
generation_count is the trap and the tooltip is blunt about it: 1 to 16 takes, each billed as its own request, and the API has no seed or variation setting, so every take is a fresh roll of the same prompt. All outputs come back as lists in the same order, with failed takes skipped unless everything failed - so your list of four clips might be three clips and a gap you can't see. That ordering guarantee is why the outputs are lists rather than one merged result, and it's why the next node runs once per clip.
Interrupting is the second one. The node stops waiting, but it sends no cancel request - a take that was already submitted may still be billed. On a 16-take run, think before you hit cancel.
And the write-it-like-a-script rule. People paste a paragraph of prose and get a monotone reading, then conclude the model is weak. It isn't: name the characters, give them lines, say who says what. A prompt that reads like a screenplay gets a scene; a prompt that reads like a caption gets a caption.
Inputs (17)
| Name | Type | Default | Description |
|---|---|---|---|
| text_prompt | STRING | Describe the voice(s), emotion, pacing, ambience, background music and sound effects, and include the lines to speak (name characters inline for dialogue). In 'audio reference' mode, refer to connected clips by order as @Audio1, @Audio2, @Audio3. A quoted line can start with a timestamp range that controls when and how long it is spoken, e.g. "[5.5s:8.0s] Wait for me!". Write the prompt in the same language as the lines to speak. Maximum 3000 characters. | |
| reference_mode | COMBO | How to condition the voice: 'text only' (describe everything in the prompt), 'audio reference' (clone up to 3 voices, tagged @Audio1-3), 'image reference' (derive a voice from one character image), or 'preset voice' (pick a built-in named voice that reads the prompt). | |
| sample_rate | COMBO | 24000 | Output sample rate in Hz. |
| speech_rate | INT | 0-50–100 | Speaking speed. 0 = normal, 100 = 2.0x, -50 = 0.5x. |
| loudness_rate | INT | 0-50–100 | Loudness. 0 = normal, 100 = 2.0x, -50 = 0.5x. |
| pitch_rate | INT | 0-12–12 | Pitch shift in semitones (-12 to 12). |
| seed | INT | 420–2147483647 | Seed controls whether the node should re-run; results are non-deterministic regardless of seed. |
| modelopt | COMBO | seed-audio-1.0 | seed-audio-1.0: 20 languages (English, Chinese, Japanese, Korean, Mexican & Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish) plus per-sentence timing control via "[5.5s:8.0s] ..." timestamps. |
| audio_formatopt | COMBO | wav | Format requested from the API (the output is decoded to AUDIO either way). |
| enable_subtitleopt | BOOLEAN | false | Return sentence and word timestamps (subtitles_json and srt outputs). |
| aigc_watermarkopt | BOOLEAN | false | Add the audible AI-generated marker at the end of the audio. |
| aigc_metadataopt | BOOLEAN | false | Implicit watermark: write AI-generation metadata into the audio header. |
| content_produceropt | STRING | Implicit watermark: name or code of the synthesis provider. | |
| produce_idopt | STRING | Implicit watermark: content production ID. | |
| content_propagatoropt | STRING | Implicit watermark: name or code of the distributor. | |
| propagate_idopt | STRING | Implicit watermark: content distribution ID. | |
| generation_countopt | INT | 11–16 | Number of separate generations to run in parallel, each billed as its own request. The API has no seed or variation setting, so every run is a fresh take of the same prompt. All outputs are lists in the same order: the next node runs once per clip. |
Outputs (5)
| Name | Type | Description |
|---|---|---|
| AUDIO | AUDIO | — |
| subtitles_json | STRING | — |
| srt | STRING | — |
| duration | FLOAT | — |
| url | STRING | — |