Nodes/ComfyUI-BytePlus-ModelArk/BytePlus Seed Voice Clone
ComfyUI Node

BytePlus Seed Voice Clone

15 seconds of audio, and a slot you pay for

By byteplus-sa·Created 8 days ago·Updated about 8 hours ago· 3
BytePlus Seed Voice Clone
  • audio
  • speaker_id
  • demo_audio
  • status_json
◄speaker_id►
◄languageen►
◄reference_text►
◄demo_text►
◄disable_volume_normalizationfalse►

Cloning a voice locally is a solved problem - F5-TTS and friends will do it zero-shot from a short clip, for free. So why would you pay for this? Three reasons, in order of how often they come up: you want the clone to work with a commercial TTS voice's prosody and language handling, you want it usable from any machine that has your key rather than the one with the model on it, or you're already producing in Seed Speech and want everything in one pipeline with subtitles to match.

This node trains a voice and gives you a speaker_id. It does not speak. You take that ID to Seed Speech TTS or into a Seed Audio reference slot.

How it works

You supply a reference clip - audio, 10 to 15 seconds of clear speech, up to 10 MB - plus the language spoken in it (default en), and the node trains into a voice slot. reference_text is optional but genuinely helps: it's the text read in the clip, and the tooltip warns that training fails if the audio differs too much from it. Note the constraint above the language field: cross-language cloning isn't supported, so an English reference gives you an English voice.

Two judgement calls: disable_volume_normalization keeps the reference clip's loudness instead of normalizing it, which the tooltip says gives closer similarity - useful when your clip has a distinctive delivery. And demo_text (4 to 80 characters, same language) makes the node return a demo clip in demo_audio so you can hear the result without a second node.

Outputs are speaker_id, demo_audio and status_json.

Where does the training go? speaker_id here is the input: a voice slot ID (S_…) bought in the Seed Speech console, or your own postpaid custom voice ID. That's the part people miss. You're not training a new voice out of thin air; you're training into a slot you've paid for.

Then, to use it: connect speaker_id to Seed Speech TTS with model seed-icl-2.0 (it lands in custom_speaker_id), or reference it from a Seed Audio audio-reference slot.

Install and the key

cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt

Restart (ComfyUI 0.31.0+), or install from Manager by searching BytePlus ModelArk. This is a Seed Speech node, so it needs the Seed Speech key - separate from the ModelArk key, saved in Settings → BytePlus or as BYTEPLUS_SEED_SPEECH_API_KEY in user/.env, ap-southeast-1 only.

Where people get burned

The slot economics, and they have teeth: each slot can be trained 15 times. Not fifteen voices - fifteen trainings. And the node trains again on every run where the inputs changed, which is the whole point of ComfyUI's cache but a nasty interaction here: swap the reference clip, tweak reference_text, change disable_volume_normalization, and you've spent another training. Fine-tune the settings on a throwaway slot, then commit.

Then the billing detail nobody reads until the invoice: the first TTS call with a cloned voice starts the slot's billing. Training isn't the trigger; speaking is. So you can train something, decide it's wrong, and pay nothing - but the moment you use it, the meter on that slot is live.

Third: a 10-to-15-second reference is short, and people submit 40 seconds and get rejected or a bad clone. Clean speech, one speaker, no music bed. The model is only as good as the sample, and this is the one node in the pack where "garbage in, garbage out" is a billing event rather than just a disappointment.

If your need is a one-off clone for a personal project, the local stack will do it for free and you already know that. Reach for this when the clone has to live in the Seed Speech ecosystem - with TTS 2.0 style control, subtitle timings, and a pipeline that doesn't need your GPU.

CategoryBytePlus ModelArk/Speech

Inputs (6)

NameTypeDefaultDescription
speaker_idSTRINGVoice slot ID (S_...) bought in the Seed Speech console, or your own postpaid custom voice ID. Each slot can be trained 15 times; every run with changed inputs trains again.
languageCOMBOenLanguage spoken in the reference clip. Cross-language cloning is not supported.
reference_textSTRINGOptional: the text read in the clip. Training fails if the audio differs too much.
demo_textSTRINGOptional: text for the demo clip (4-80 characters, same language).
disable_volume_normalizationBOOLEANfalseKeep the reference clip's loudness instead of normalizing it (closer similarity).
audioAUDIOReference voice clip, e.g. from Load Audio: 10-15 s of clear speech, up to 10 MB.

Outputs (3)

NameTypeDescription
speaker_idSTRING—
demo_audioAUDIO—
status_jsonSTRING—