BytePlus Seed Speech TTS
A voiceover node with an emotion dial
- audio
- subtitles_json
- srt
Local TTS is one of the genuinely solved parts of this ecosystem - Kokoro on a CPU, Chatterbox under 8GB, F5-TTS when you want speed. So the case for a hosted TTS node isn't "you can't do this locally." It's specific: you want a particular commercial voice, you want style instruction in plain language, or you want cloned voices without training anything. Seed Speech TTS does all three, and the style-instruction field is the part local models still struggle to match.
And speech is cheap. Compared to video seconds, this is the node in the pack that won't hurt to use.
How it works
Text in, audio out, plus optional subtitles. Three model tiers decide what "voice" means:
seed-tts-2.0- the modern voice list.voicepicks one of the named presets (the default isen_female_stokie_uranus_bigtts). Each voice supports specific languages, so check the console's voice list if you need something non-English.seed-tts-1.0- older speaker IDs. These go incustom_speaker_id, which overridesvoice.seed-icl-2.0/seed-icl-1.0- cloned voices. Connect thespeaker_idoutput of BytePlus Seed Voice Clone here (or type a TTS 1.0 speaker ID).
The fields worth your attention
textis what gets spoken. Everything else is how.context_textis the interesting one, and it's 2.0-only: a style instruction or preceding dialogue - "Speak slowly, sounding heartbroken." It isn't billed, and it's the closest thing here to directing a performance.emotionplusemotion_scale(1–5) for voices that support it; empty means neutral.speech_rate,loudness_rate(−50 to 100, where 100 is 2x) andpitch(±12 semitones) are the delivery controls. Speaker IDs from a clone behave differently from presets here - cloned voices gettone_fidelityas an option, which keeps them as close as possible to the training audio's voice, emotion and accent, same language only.enable_subtitlereturns timestamps. On 2.0 voices that's sentence subtitles, on 1.0 voices word timestamps - so choose the model with that in mind if you need SRT.sample_rate(2.0 does 24000/16000/8000),explicit_language(pin it to read only that language;autohandles mixed Chinese and English),detect_language, andsilence_durationfor a trailing pause.- The text-hygiene switches:
filter_markdownstrips syntax so**bold**is read as "bold",read_emojistops emoji being silently dropped,read_latexreads formulas aloud (and turns on markdown filtering),read_parenthesesmakes bracketed asides audible instead of skipped.unsupported_char_ratiodecides when the node should give up rather than mangle mostly-foreign text. use_cachereuses audio synthesized from identical text within the last hour - no timestamps from cache, so skip it when you need subtitles.seedisn't sent to the API at all; it exists only to make the node synthesize again.
Outputs: audio (wire into Save Audio, an ASR node, or Seedance's reference audio), subtitles_json, and srt.
Install and the key that trips everyone
cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
Restart (ComfyUI 0.31.0+), or install via Manager by searching BytePlus ModelArk. Speech nodes use a different key from ModelArk: create one in the Seed Speech console (activate the services first) and save it in Settings → BytePlus as the Seed Speech key, or as BYTEPLUS_SEED_SPEECH_API_KEY in user/.env. Seed Speech runs in ap-southeast-1 only.
Where people get burned
Pasting the ModelArk key into the speech row. You'll get Invalid X-Api-Key - not the friendly 401 you get elsewhere - because these are separate products with separate keys. Nothing is wrong with your account; you're holding the wrong credential.
Second: voices and languages aren't freely mixable. A 2.0 voice supports a defined set of languages, so the fastest path to Japanese isn't explicit_language = ja on an English voice, it's picking a voice that does Japanese. Third, custom_speaker_id silently wins over voice, which is exactly right the first time you connect a clone and exactly confusing three weeks later when you've forgotten you did.
Inputs (24)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | seed-tts-2.0 | seed-tts-2.0 for the voice list; seed-icl-* for cloned voices (custom_speaker_id). |
| text | STRING | — | |
| voice | COMBO | en_female_stokie_uranus_bigtts | TTS 2.0 voice. Each voice supports specific languages; see the Seed Speech voice list. |
| custom_speaker_id | STRING | Overrides voice: a cloned voice (connect Seed Voice Clone's speaker_id, model seed-icl-2.0) or a TTS 1.0 speaker ID. | |
| context_text | STRING | TTS 2.0 only: style instruction or preceding dialogue, e.g. 'Speak slowly, sounding heartbroken.' Not billed. | |
| emotion | STRING | Emotion for voices that support it, e.g. happy, sad, angry. Empty = neutral. | |
| emotion_scale | INT | 41–5 | — |
| speech_rate | INT | 0-50–100 | 100 = 2x speed, -50 = 0.5x. |
| loudness_rate | INT | 0-50–100 | 100 = 2x volume, -50 = 0.5x. |
| pitch | INT | 0-12–12 | Pitch shift in semitones. |
| sample_rate | COMBO | 24000 | Hz. seed-tts-2.0 supports 24000, 16000 and 8000. |
| explicit_language | COMBO | auto | Read only text in this language. auto handles mixed Chinese and English. |
| silence_duration | INT | 00–30000 | Silence added after the last sentence, in ms. |
| filter_markdown | BOOLEAN | false | Strip Markdown syntax so **bold** is read as 'bold'. |
| enable_subtitle | BOOLEAN | false | Return timestamps (Chinese and English): subtitles for 2.0 voices, word timestamps for 1.0 voices. |
| detect_language | BOOLEAN | false | Detect the text language automatically. |
| context_language | COMBO | default | Reference language for Western European text: default English, id Indonesian, es Mexican Spanish, pt Brazilian Portuguese. |
| read_emoji | BOOLEAN | false | Keep emoji in the text instead of filtering them out. |
| read_latex | BOOLEAN | false | Read LaTeX formulas aloud (also enables filter_markdown). |
| read_parentheses | BOOLEAN | false | Read text inside parentheses (filtered out by default). |
| unsupported_char_ratio | FLOAT | 0.300–1 | Fail when more than this share of the text is in an unsupported language. |
| use_cache | BOOLEAN | false | Reuse audio synthesized for identical text in the last hour (no timestamps from cache). |
| tone_fidelity | BOOLEAN | false | seed-icl-2.0 only: stay as close as possible to the training audio's voice, emotion and accent (same language only). |
| seed | INT | 00–18446744073709550000 | Not sent to the API; change it to synthesize again. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| audio | AUDIO | — |
| subtitles_json | STRING | — |
| srt | STRING | — |