Nodes/ComfyUI-kaola-ace-step/ACE-Step Understand
ComfyUI Node

ACE-Step Understand

A ComfyUI node in Audio/ACE-Step with 12 inputs and 6 outputs.

By kana112233·Created 6 months ago·Updated 6 months ago· 27
ACE-Step Understand
  • audio
  • analysis_text
  • caption
  • duration
  • bpm
  • keyscale
  • lyrics
checkpoint_dirAce-Step1.5
config_pathacestep-v15-turbo
lm_model_pathacestep-5Hz-lm-1.7B
target_duration30.00
deviceauto
languageauto
temperature0.3
top_k0
top_p0.90
repetition_penalty1.0
thinkingtrue
CategoryAudio/ACE-Step

Inputs (12)

NameTypeDefaultDescription
audioAUDIOThe audio signal to be analyzed.
checkpoint_dirCOMBOAce-Step1.5Directory containing ACE-Step model weights (DiT model).
config_pathCOMBOacestep-v15-turboSpecific model configuration to use (e.g., v1.5 turbo).
lm_model_pathCOMBOacestep-5Hz-lm-1.7BPath to the language model used for audio analysis and understanding.
target_durationFLOAT30.0010–600Target duration to reference during analysis.
deviceCOMBOautoComputing platform to run the model on.
languageoptCOMBOautoHint the model about the vocal language in the audio.
temperatureoptFLOAT0.30–2Sampling temperature. Lower (0.0-0.3) = more precise/faithful, Higher (0.5+) = more creative. Try 0.1 for better accuracy.
top_koptINT00–100Top-K sampling. 0 = disabled. Lower values (e.g., 20-50) can improve accuracy by limiting token choices.
top_poptFLOAT0.900–1Top-P (nucleus) sampling. 1.0 = disabled. Lower values (e.g., 0.8-0.9) can improve accuracy.
repetition_penaltyoptFLOAT1.00.5–2Repetition penalty. 1.0 = no penalty. Higher values (1.1-1.3) reduce repetitive lyrics.
thinkingoptBOOLEANtrueWhether to show the language model's Chain-of-Thought reasoning.

Outputs (6)

NameTypeDescription
analysis_textSTRING
captionSTRING
durationFLOAT
bpmSTRING
keyscaleSTRING
lyricsSTRING