Nodes/ComfyUI-API-Toolkit/ElevenLabs - Voice Design
ComfyUI Node

ElevenLabs - Voice Design

Invent a voice from a description — no recording required

By IxMxAMAR·Created 6 months ago·Updated 3 days ago· 1
ElevenLabs - Voice Design
  • reference_audio
  • preview_audio
  • generated_voice_id
◄api_key►
◄textHello! This is a preview of the designed voice.►
◄voice_description►
◄modeleleven_multilingual_ttv_v2►
◄auto_generate_textfalse►
◄loudness0.50►
◄guidance_scale5.0►
◄seed0►
◄should_enhancefalse►
◄prompt_strength0.50►

Voice cloning needs a recording of the person you want to sound like. Voice design needs nothing but a sentence. This node takes a written description - age, gender, accent, tone, register - and returns a brand-new voice that matches it, complete with an audio preview of what that voice sounds like. No mic, no reference clip, no person involved.

Three required inputs:

  • api_key.
  • text - the sample line the preview will read. Defaults to "Hello! This is a preview of the designed voice." Change it to something relevant to your project so you audition the voice on real copy.
  • voice_description - the actual prompt. This is the important one. "Warm, mid-aged British woman, soft-spoken, slightly husky, like a friendly audiobook narrator." The richer the description, the closer the result. Age, gender, accent, and tone are the four axes that matter most.

Two outputs:

  • preview_audio - the sample line read in the designed voice. Listen before you commit.
  • generated_voice_id - the temporary ID of this design, which is the handoff to the next node.

The key fact: this node does not save the voice. The preview is ephemeral and generated_voice_id is a temporary token. If you like what you hear, you feed that ID into AIS_EL_VoiceCreate, which permanently saves it to your ElevenLabs library and returns a real voice_id you can then use in TTS forever. Design and save are deliberately two steps, so you don't clutter your library with every sketch.

Why you'd reach for it

The killer use is character work: you need a voice that doesn't exist yet - an old pirate, a cheerful robot, a stuffy butler - and there's no one to record. Design gets you from description to audio in one shot, and you can iterate: tweak the description, re-run, listen, until the character sounds right. It also beats scrolling the premade voice list when you need something specific the catalog doesn't have.

Installing it

Part of the ComfyUI API Toolkit pack. Manager: search "API Toolkit". Manual:

cd ComfyUI/custom_nodes
git clone https://github.com/IxMxAMAR/ComfyUI-API-Toolkit
cd ComfyUI-API-Toolkit
pip install -r requirements.txt

Restart. Needs requests and soundfile.

Gotchas

  • The generated_voice_id is temporary. If you want to keep the voice, chain into VoiceCreate before you move on - the design token doesn't survive forever.
  • Design is a preview API call, so it's still metered, and iterating means several calls. Keep descriptions tight and audition on short sample text.
  • The two-step split trips people who expect one node to "make a voice." It doesn't. Preview here, save there, that's the design.
CategoryAPI Toolkit/ElevenLabs/Voice

Inputs (11)

NameTypeDefaultDescription
api_keySTRING—
textSTRINGHello! This is a preview of the designed voice.Sample text, 100-1000 characters. Shorter text is replaced by auto-generated sample text.
voice_descriptionSTRINGDescribe the voice you want: age, gender, accent, tone, etc. 20-1000 characters.
modeloptCOMBOeleven_multilingual_ttv_v2Voice design model. eleven_ttv_v3 is the only one that accepts reference_audio.
auto_generate_textoptBOOLEANfalseGenerate the sample text from the voice description instead of using `text`.
loudnessoptFLOAT0.50-1–1Volume of the generated voice. -1 = quietest, 1 = loudest, 0 is roughly -24 LUFS.
guidance_scaleoptFLOAT5.00–100How closely the voice follows the description. High values can sound robotic.
seedoptINT00–2147483647Same seed with the same inputs produces the same voice. 0 = random.
should_enhanceoptBOOLEANfalseExpand a short description into a more detailed one before generating.
reference_audiooptAUDIOReference voice to design from. Only supported with eleven_ttv_v3.
prompt_strengthoptFLOAT0.500–1Balance of description vs reference_audio: 0 = almost no description influence, 1 = almost no reference influence. Used only with reference_audio.

Outputs (2)

NameTypeDescription
preview_audioAUDIO—
generated_voice_idSTRING—