Nodes/ComfyUI-OmniVoice_CRT/OmniVoice Generate Audio
ComfyUI Node

OmniVoice Generate Audio

A ComfyUI node in OmniVoice/Process with 18 inputs and 2 outputs.

By PGCRT·Created 4 months ago·Updated 4 months ago· 3
OmniVoice Generate Audio
  • pipe
  • audio
  • status
textHello from OmniVoice.
languageauto
style_gendernone
style_agenone
style_pitchnone
style_accentnone
num_step32
guidance_scale2.0
t_shift0.10
layer_penalty_factor5.0
position_temperature5.0
class_temperature0.00
speed1.00
seed0
use_durationfalse
duration10.0
postprocess_outputtrue
CategoryOmniVoice/Process

Inputs (18)

NameTypeDefaultDescription
pipeOMNIVOICE_PIPEPipe output from OmniVoice Load Model
textSTRINGHello from OmniVoice.Target speech text
languageCOMBOautoLanguage hint for TTS; auto lets model infer language
style_genderCOMBOnoneno ref audio only
style_ageCOMBOnoneno ref audio only
style_pitchCOMBOnoneno ref audio only
style_accentCOMBOnoneno ref audio only
num_stepINT324–128Diffusion sampling steps (higher = slower, often cleaner)
guidance_scaleFLOAT2.00–20Classifier-free guidance strength
t_shiftFLOAT0.100–1Diffusion timestep shift; lower tends to favor low-SNR detail
layer_penalty_factorFLOAT5.00–20Penalty encouraging earlier codebook layers to unmask first
position_temperatureFLOAT5.00–20Temperature for position selection during generation
class_temperatureFLOAT0.000–2Token class sampling temperature (0 = greedy)
speedFLOAT1.000.25–4
seedINT00–2147483647Random seed for reproducibility
use_durationBOOLEANfalseEnable fixed output duration override
durationFLOAT10.00–120Target output duration in seconds (used only when enabled)
postprocess_outputBOOLEANtrueApply output postprocessing (trim/fade/pad cleanup)

Outputs (2)

NameTypeDescription
audioAUDIO
statusSTRING