Nodes/ComfyUI-BytePlus-ModelArk/BytePlus Seed Audio 1.0
ComfyUI Node

BytePlus Seed Audio 1.0

Two minutes of dialogue, music and foley from one prompt

By byteplus-sa·Created 8 days ago·Updated about 8 hours ago· 3
BytePlus Seed Audio 1.0
    • AUDIO
    • subtitles_json
    • srt
    • duration
    • url
    ◄text_prompt►
    ◄reference_mode▾►
    ◄sample_rate24000►
    ◄speech_rate0►
    ◄loudness_rate0►
    ◄pitch_rate0►
    ◄seed42►
    ◄modelseed-audio-1.0►
    ◄audio_formatwav►
    ◄enable_subtitlefalse►
    ◄aigc_watermarkfalse►
    ◄aigc_metadatafalse►
    ◄content_producer►
    ◄produce_id►
    ◄content_propagator►
    ◄propagate_id►
    ◄generation_count1►

    Most TTS nodes read text. This one stages a scene. You describe the voices, the emotion, the pacing, the room, the background music and the sound effects - and then you write the lines, naming characters inline - and you get back up to two minutes of produced audio. Dialogue with two speakers. A voiceover with a music bed and a door slam in the right place. It's the one node in the pack where the prompt is doing sound design rather than synthesis.

    The pack's Sound Design and Podcast Clip templates both lean on it, and that's the honest framing: it's for scenes, not for reading a paragraph of documentation aloud. If you just need narration, use Seed Speech TTS - it's more controllable, line by line.

    How it works

    text_prompt carries everything, up to 3000 characters, and it's worth writing like a script rather than a sentence. Name your characters, put their lines in the prompt, and describe the environment around them. Write it in the same language as the lines you want spoken.

    reference_mode is the structural choice, and it changes what else appears on the node:

    • text only - everything comes from the prompt. No references, no @Audio tags.
    • audio reference - up to three connected clips (reference_audio_1 … _3, each up to 30 s), or a speaker ID / cloned voice ID / audio URL in the matching ref_audio_N_source fields. You refer to them in the prompt as @Audio1, @Audio2, @Audio3. Slots must be filled from 1 upward; gaps are errors.
    • image reference - one character image (reference_image, or ref_image_url as a link), and the model derives a voice from it. Can't be combined with reference audio.
    • preset voice - pick a built-in TTS 2.0 voice from preset_voice and it reads the prompt. No clone, no tags.

    Then the mix controls: sample_rate, speech_rate, loudness_rate (−50 to 100, 100 = 2x), pitch_rate (±12 semitones), and seed - which, as everywhere in this pack, only decides whether the node re-runs. model defaults to seed-audio-1.0 and covers 20 languages (English, Chinese, Japanese, Korean, both Spanishes, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish).

    The timestamp syntax is the feature that makes this useful for video: a quoted line can start with a range, like [5.5s:8.0s] Wait for me!, and that controls when it's spoken and how long it takes. Synchronising a line to a shot becomes arithmetic instead of trial and error.

    Outputs are AUDIO, subtitles_json, srt, duration and url - and every one of them is a list, because of generation_count. Set enable_subtitle if you want the timestamps; audio_format picks what's requested from the API (WAV, MP3, OGG Opus, PCM), though it's decoded to AUDIO either way.

    One field group worth knowing about even if you never touch it: aigc_watermark adds the audible AI marker, and aigc_metadata plus content_producer / produce_id / content_propagator / propagate_id write provenance into the audio header. If you're producing something that has to declare itself, that's where.

    Install and the key

    cd ComfyUI/custom_nodes
    git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
    pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
    

    Restart (ComfyUI 0.31.0+), or install through Manager by searching BytePlus ModelArk. Speech is a separate product with a separate key: create it in the Seed Speech console (activate the services first) and save it as the Seed Speech key in Settings → BytePlus, or BYTEPLUS_SEED_SPEECH_API_KEY in user/.env. Seed Speech is ap-southeast-1 only.

    Where people get burned

    generation_count is the trap and the tooltip is blunt about it: 1 to 16 takes, each billed as its own request, and the API has no seed or variation setting, so every take is a fresh roll of the same prompt. All outputs come back as lists in the same order, with failed takes skipped unless everything failed - so your list of four clips might be three clips and a gap you can't see. That ordering guarantee is why the outputs are lists rather than one merged result, and it's why the next node runs once per clip.

    Interrupting is the second one. The node stops waiting, but it sends no cancel request - a take that was already submitted may still be billed. On a 16-take run, think before you hit cancel.

    And the write-it-like-a-script rule. People paste a paragraph of prose and get a monotone reading, then conclude the model is weak. It isn't: name the characters, give them lines, say who says what. A prompt that reads like a screenplay gets a scene; a prompt that reads like a caption gets a caption.

    CategoryBytePlus ModelArk/Speech

    Inputs (17)

    NameTypeDefaultDescription
    text_promptSTRINGDescribe the voice(s), emotion, pacing, ambience, background music and sound effects, and include the lines to speak (name characters inline for dialogue). In 'audio reference' mode, refer to connected clips by order as @Audio1, @Audio2, @Audio3. A quoted line can start with a timestamp range that controls when and how long it is spoken, e.g. "[5.5s:8.0s] Wait for me!". Write the prompt in the same language as the lines to speak. Maximum 3000 characters.
    reference_modeCOMBOHow to condition the voice: 'text only' (describe everything in the prompt), 'audio reference' (clone up to 3 voices, tagged @Audio1-3), 'image reference' (derive a voice from one character image), or 'preset voice' (pick a built-in named voice that reads the prompt).
    sample_rateCOMBO24000Output sample rate in Hz.
    speech_rateINT0-50–100Speaking speed. 0 = normal, 100 = 2.0x, -50 = 0.5x.
    loudness_rateINT0-50–100Loudness. 0 = normal, 100 = 2.0x, -50 = 0.5x.
    pitch_rateINT0-12–12Pitch shift in semitones (-12 to 12).
    seedINT420–2147483647Seed controls whether the node should re-run; results are non-deterministic regardless of seed.
    modeloptCOMBOseed-audio-1.0seed-audio-1.0: 20 languages (English, Chinese, Japanese, Korean, Mexican & Castilian Spanish, Indonesian, German, Brazilian Portuguese, French, Thai, Vietnamese, Malay, Filipino, Italian, Russian, Dutch, Polish, Turkish, Swedish) plus per-sentence timing control via "[5.5s:8.0s] ..." timestamps.
    audio_formatoptCOMBOwavFormat requested from the API (the output is decoded to AUDIO either way).
    enable_subtitleoptBOOLEANfalseReturn sentence and word timestamps (subtitles_json and srt outputs).
    aigc_watermarkoptBOOLEANfalseAdd the audible AI-generated marker at the end of the audio.
    aigc_metadataoptBOOLEANfalseImplicit watermark: write AI-generation metadata into the audio header.
    content_produceroptSTRINGImplicit watermark: name or code of the synthesis provider.
    produce_idoptSTRINGImplicit watermark: content production ID.
    content_propagatoroptSTRINGImplicit watermark: name or code of the distributor.
    propagate_idoptSTRINGImplicit watermark: content distribution ID.
    generation_countoptINT11–16Number of separate generations to run in parallel, each billed as its own request. The API has no seed or variation setting, so every run is a fresh take of the same prompt. All outputs are lists in the same order: the next node runs once per clip.

    Outputs (5)

    NameTypeDescription
    AUDIOAUDIO—
    subtitles_jsonSTRING—
    srtSTRING—
    durationFLOAT—
    urlSTRING—