Nodes/ComfyUI-GGUF-Loader/Scenema Audio Generate ⚡
ComfyUI Node

Scenema Audio Generate ⚡

Expressive text-to-speech with zero-shot voice cloning on the Scenema audio diffusion model.

By ChrisColeTech·Created 13 days ago·Updated about 3 hours ago· 6
Scenema Audio Generate ⚡
  • model
  • clip
  • vae
  • ref_latent
  • identity_reference
  • audio
presetCustom
voice_descriptionMale, late 60s. Deep, gravelly. Slow and deliberate. The weight of the cosmos in every word.
gender
speech_textLook again at that dot. That's here. That's home. That's us.
scene
seed0
custom_scene
action_tags
languageEnglish
pace1.5
skip_vcfalse
vc_steps25
vc_cfg_rate0.50
Category🤖 CCTech/Scenema

Inputs (18)

NameTypeDefaultDescription
modelMODEL
clipCLIP
vaeVAE
presetCOMBOCustomApply a preset's voice, gender, scene, and performance tags. The speech_text field is always used as entered.
voice_descriptionSTRINGMale, late 60s. Deep, gravelly. Slow and deliberate. The weight of the cosmos in every word.Describe the voice: age, gender presentation, timbre, accent, delivery style.
genderCOMBOGrammatical gender used for pronouns in the compiled prompt.
speech_textSTRINGLook again at that dot. That's here. That's home. That's us.The text to speak. Use [bracketed cues] inline for mid-speech performance direction: [He laughs], [She whispers]. Long text is auto-split at sentence boundaries.
sceneCOMBOAcoustic environment injected into the prompt.
seedINT00–18446744073709550000
custom_sceneoptSTRINGFreeform scene description. When non-empty, overrides the scene dropdown.
action_tagsoptSTRINGDelivery cues, one per line. Each becomes a stage direction the model performs.
languageoptCOMBOEnglishTarget language. Write speech_text in that language.
paceoptFLOAT1.50.5–3Duration budget multiplier. Higher = slower speech. 1.5 is the validated default.
ref_latentoptLATENTOptional voice reference (Scenema VAE Encode output) for zero-shot cloning.
skip_vcoptBOOLEANfalseSkip the final SeedVC identity-consistency pass.
identity_referenceoptAUDIOOptional fixed voice identity for SeedVC. This is separate from the LTX A2V ref_latent input.
vc_stepsoptINT251–200
vc_cfg_rateoptFLOAT0.500–2

Outputs (1)

NameTypeDescription
audioAUDIO