Nodes/ComfyUI-BytePlus-ModelArk/BytePlus Seed Speech TTS
ComfyUI Node

BytePlus Seed Speech TTS

A voiceover node with an emotion dial

By byteplus-sa·Created 8 days ago·Updated about 8 hours ago· 3
BytePlus Seed Speech TTS
    • audio
    • subtitles_json
    • srt
    ◄modelseed-tts-2.0►
    ◄text►
    ◄voiceen_female_stokie_uranus_bigtts►
    ◄custom_speaker_id►
    ◄context_text►
    ◄emotion►
    ◄emotion_scale4►
    ◄speech_rate0►
    ◄loudness_rate0►
    ◄pitch0►
    ◄sample_rate24000►
    ◄explicit_languageauto►
    ◄silence_duration0►
    ◄filter_markdownfalse►
    ◄enable_subtitlefalse►
    ◄detect_languagefalse►
    ◄context_languagedefault►
    ◄read_emojifalse►
    ◄read_latexfalse►
    ◄read_parenthesesfalse►
    ◄unsupported_char_ratio0.30►
    ◄use_cachefalse►
    ◄tone_fidelityfalse►
    ◄seed0►

    Local TTS is one of the genuinely solved parts of this ecosystem - Kokoro on a CPU, Chatterbox under 8GB, F5-TTS when you want speed. So the case for a hosted TTS node isn't "you can't do this locally." It's specific: you want a particular commercial voice, you want style instruction in plain language, or you want cloned voices without training anything. Seed Speech TTS does all three, and the style-instruction field is the part local models still struggle to match.

    And speech is cheap. Compared to video seconds, this is the node in the pack that won't hurt to use.

    How it works

    Text in, audio out, plus optional subtitles. Three model tiers decide what "voice" means:

    • seed-tts-2.0 - the modern voice list. voice picks one of the named presets (the default is en_female_stokie_uranus_bigtts). Each voice supports specific languages, so check the console's voice list if you need something non-English.
    • seed-tts-1.0 - older speaker IDs. These go in custom_speaker_id, which overrides voice.
    • seed-icl-2.0 / seed-icl-1.0 - cloned voices. Connect the speaker_id output of BytePlus Seed Voice Clone here (or type a TTS 1.0 speaker ID).

    The fields worth your attention

    • text is what gets spoken. Everything else is how.
    • context_text is the interesting one, and it's 2.0-only: a style instruction or preceding dialogue - "Speak slowly, sounding heartbroken." It isn't billed, and it's the closest thing here to directing a performance.
    • emotion plus emotion_scale (1–5) for voices that support it; empty means neutral. speech_rate, loudness_rate (−50 to 100, where 100 is 2x) and pitch (±12 semitones) are the delivery controls. Speaker IDs from a clone behave differently from presets here - cloned voices get tone_fidelity as an option, which keeps them as close as possible to the training audio's voice, emotion and accent, same language only.
    • enable_subtitle returns timestamps. On 2.0 voices that's sentence subtitles, on 1.0 voices word timestamps - so choose the model with that in mind if you need SRT.
    • sample_rate (2.0 does 24000/16000/8000), explicit_language (pin it to read only that language; auto handles mixed Chinese and English), detect_language, and silence_duration for a trailing pause.
    • The text-hygiene switches: filter_markdown strips syntax so **bold** is read as "bold", read_emoji stops emoji being silently dropped, read_latex reads formulas aloud (and turns on markdown filtering), read_parentheses makes bracketed asides audible instead of skipped. unsupported_char_ratio decides when the node should give up rather than mangle mostly-foreign text.
    • use_cache reuses audio synthesized from identical text within the last hour - no timestamps from cache, so skip it when you need subtitles. seed isn't sent to the API at all; it exists only to make the node synthesize again.

    Outputs: audio (wire into Save Audio, an ASR node, or Seedance's reference audio), subtitles_json, and srt.

    Install and the key that trips everyone

    cd ComfyUI/custom_nodes
    git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
    pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
    

    Restart (ComfyUI 0.31.0+), or install via Manager by searching BytePlus ModelArk. Speech nodes use a different key from ModelArk: create one in the Seed Speech console (activate the services first) and save it in Settings → BytePlus as the Seed Speech key, or as BYTEPLUS_SEED_SPEECH_API_KEY in user/.env. Seed Speech runs in ap-southeast-1 only.

    Where people get burned

    Pasting the ModelArk key into the speech row. You'll get Invalid X-Api-Key - not the friendly 401 you get elsewhere - because these are separate products with separate keys. Nothing is wrong with your account; you're holding the wrong credential.

    Second: voices and languages aren't freely mixable. A 2.0 voice supports a defined set of languages, so the fastest path to Japanese isn't explicit_language = ja on an English voice, it's picking a voice that does Japanese. Third, custom_speaker_id silently wins over voice, which is exactly right the first time you connect a clone and exactly confusing three weeks later when you've forgotten you did.

    CategoryBytePlus ModelArk/Speech

    Inputs (24)

    NameTypeDefaultDescription
    modelCOMBOseed-tts-2.0seed-tts-2.0 for the voice list; seed-icl-* for cloned voices (custom_speaker_id).
    textSTRING—
    voiceCOMBOen_female_stokie_uranus_bigttsTTS 2.0 voice. Each voice supports specific languages; see the Seed Speech voice list.
    custom_speaker_idSTRINGOverrides voice: a cloned voice (connect Seed Voice Clone's speaker_id, model seed-icl-2.0) or a TTS 1.0 speaker ID.
    context_textSTRINGTTS 2.0 only: style instruction or preceding dialogue, e.g. 'Speak slowly, sounding heartbroken.' Not billed.
    emotionSTRINGEmotion for voices that support it, e.g. happy, sad, angry. Empty = neutral.
    emotion_scaleINT41–5—
    speech_rateINT0-50–100100 = 2x speed, -50 = 0.5x.
    loudness_rateINT0-50–100100 = 2x volume, -50 = 0.5x.
    pitchINT0-12–12Pitch shift in semitones.
    sample_rateCOMBO24000Hz. seed-tts-2.0 supports 24000, 16000 and 8000.
    explicit_languageCOMBOautoRead only text in this language. auto handles mixed Chinese and English.
    silence_durationINT00–30000Silence added after the last sentence, in ms.
    filter_markdownBOOLEANfalseStrip Markdown syntax so **bold** is read as 'bold'.
    enable_subtitleBOOLEANfalseReturn timestamps (Chinese and English): subtitles for 2.0 voices, word timestamps for 1.0 voices.
    detect_languageBOOLEANfalseDetect the text language automatically.
    context_languageCOMBOdefaultReference language for Western European text: default English, id Indonesian, es Mexican Spanish, pt Brazilian Portuguese.
    read_emojiBOOLEANfalseKeep emoji in the text instead of filtering them out.
    read_latexBOOLEANfalseRead LaTeX formulas aloud (also enables filter_markdown).
    read_parenthesesBOOLEANfalseRead text inside parentheses (filtered out by default).
    unsupported_char_ratioFLOAT0.300–1Fail when more than this share of the text is in an unsupported language.
    use_cacheBOOLEANfalseReuse audio synthesized for identical text in the last hour (no timestamps from cache).
    tone_fidelityBOOLEANfalseseed-icl-2.0 only: stay as close as possible to the training audio's voice, emotion and accent (same language only).
    seedINT00–18446744073709550000Not sent to the API; change it to synthesize again.

    Outputs (3)

    NameTypeDescription
    audioAUDIO—
    subtitles_jsonSTRING—
    srtSTRING—