Nodes/ComfyUI-AuK_Doc/AuK Generate / Edit_Doc
ComfyUI Node

AuK Generate / Edit_Doc

TTS, voice cloning and speech edits from one instruction box

By DocWorkBox·Created 22 days ago·Updated 22 days ago· 4
AuK Generate / Edit_Doc
  • engine
  • input_audio
  • llm_config
  • generated audio
  • model instruction
  • Prompt Enhancer info
◄instructionGenerate speech based on the following description: "a calm, warm female voice". The content to speak is: "Hello, welcome to AuK.".►
◄generation_seconds0.0►
◄use_prompt_enhancertrue►
◄seed42►
◄nfe_steps32►
◄cfg_strength2.0►
◄sway_sampling_coef-1.0►

Every TTS node in ComfyUI answers the same question: what should this voice say? AuK answers a different one. Its instruction box lets you say "keep the voice, change the words," "say this faster," "strip the music off this track," "make it sound less accented" - and get the same speaker back with one thing changed. That's the reason this node exists, and the reason it's worth the download, because almost nothing else in the graph does it.

The job it does that a TTS node doesn't

Tencent's AuK covers instruction TTS (describe the voice, get speech), zero-shot cloning from a reference clip, content editing (replace, add, delete words), voice and emotion conversion, speed and pitch adjustment, denoising, and vocal separation. The KB's read of this layer is blunt - audio is the newest, thinnest part of the stack, bolted on once video got good enough to want a soundtrack (audio-generation.md). Chatterbox and F5-TTS own "make me a voice"; AuK owns "take this recording and hand it back edited," which is a job the community mostly does by re-generating and hoping.

ComfyUI-AuK_Doc is DocWorkBox's repackage of that model as a self-contained node pack - no separate upstream source checkout, node names suffixed _Doc so they don't collide with anything official. All four nodes in it are on the canvas, but this is the one doing work.

How it works

Two models run per call. Qwen2.5-Omni-3B is the text-and-audio encoder - it reads your instruction and any reference waveform. The AuK DiT (a flow-matching / CFM model, Base or Flash) denoises in that embedding space, and AuK's own VAE decodes to audio. Your reference audio is VAE-encoded into the same sequence as the target, which is the mechanical reason for the 30-second ceiling: source and output share one budget, not two.

The Prompt Enhancer is the part that makes free-form instructions work, and it's the same path the official Gradio demo uses: an LLM classifies your request into a task type (instruction TTS, zero-shot TTS, an edit subtype, restoration, separation), ASR transcribes your input audio if you connected one (Tencent Cloud recording ASR when credentials exist, otherwise FunASR's SenseVoiceSmall on CPU), a target duration gets estimated, and your sentence is rewritten into the official template wording. That rewritten string is the model instruction output - read it when a run surprises you.

The inputs that matter

instruction is where you live. Copy the templates out of the pack's Markdown Note (they're the official Cookbook lines) and swap the placeholders - the model is fussy about the phrasing that ships in those templates, and improvising is how you get garbage.

generation_seconds is the trap. With Enable Prompt Enhancer on (the default), leave it at 0 and let the enhancer estimate; with the enhancer off, 0 is an error, so the "why did nothing happen" answer is almost always this field.

input_audio decides which mode you're in: unconnected means instruction TTS, connected means that clip is the reference voice (zero-shot cloning) or the thing being edited. Batch size must be 1 and multichannel audio gets averaged to mono.

engine comes from the pack's loader, and llm_config is optional - leave it unwired and the enhancer reads LLM_API_KEY, LLM_BASE_URL and LLM_MODEL_NAME from the server environment instead.

The three advanced fields belong to AuK, not to you: nfe_steps, cfg_strength and sway_sampling_coef are 32 / 2 / −1 for Base. Load AuK-Flash and the node flatly refuses anything but 4 / 0 / −1, because that's the Flash recipe.

Outputs

generated audio is a normal AUDIO - wire it to Save Audio or the newer preview nodes. This is not an output node, so nothing runs until something downstream asks for audio. The other two sockets (model instruction, Prompt Enhancer info) are text and exist for debugging: the second one prints the detected task, the final target duration and the ASR transcript.

Install

ComfyUI Manager → search ComfyUI-AuK_Doc, or by hand:

cd ComfyUI/custom_nodes
git clone https://github.com/DocWorkBox/ComfyUI-AuK_Doc
cd ComfyUI-AuK_Doc
python -m pip install -r requirements.txt

Use the Python that runs ComfyUI, then restart. Requirements are the heavy part - transformers>=4.52,<5, qwen-omni-utils, funasr, modelscope, tencentcloud-sdk-python-asr, silero-vad, torchdiffeq, openai - while Torch/TorchAudio/NumPy come from your ComfyUI install unpinned. Audio I/O goes through SoundFile specifically to dodge the TorchCodec requirement in recent TorchAudio.

Then the weights: auk_base.safetensors or auk_flash.safetensors, plus vae.safetensors and config.yaml, in ComfyUI/models/auk/AuK (or AuK-Flash), and a complete Qwen2.5-Omni-3B directory in ComfyUI/models/text_encoders/. GGUF is not a substitute for the Qwen encoder - it only matters for the optional local-LLM path. Four example workflows ship in workflows/; drag one in before building your own.

Where people get burned

  • The 30-second wall. Source plus target must fit in 30 s, and the error message tells you the split ("1.20s + 5.00s"). Editing a 20-second clip and asking for 15 seconds back doesn't fit. Work in clips.
  • Enhancer on with no credentials raises a missing-LLM-config error rather than silently doing nothing. Connect a settings node or set the env vars (see the OpenAI / llama.cpp articles in this pack).
  • Lyric edits need clean vocals. Feed it a song with the instrumental in it and you've asked it to edit karaoke, not words. Separate first.
  • VRAM is shared with the rest of your graph. With the loader in low_vram or balanced, each run calls ComfyUI's unload_all_models() and empties the cache, so your checkpoints get evicted. That's intentional, not a leak.
CategoryAuK_Doc

Inputs (10)

NameTypeDefaultDescription
engineAUK_ENGINE_Doc—
instructionSTRINGGenerate speech based on the following description: "a calm, warm female voice". The content to speak is: "Hello, welcome to AuK.".Natural-language AuK request. Prompt Enhancer can convert a free-form request into the model instruction.
generation_secondsFLOAT0.00–30Generated target duration in seconds. 0 lets Prompt Enhancer estimate it. When Prompt Enhancer is disabled, enter a value above 0.
use_prompt_enhancerBOOLEANtrueUses the same PE preparation path as the AuK Gradio demo. Credentials are read from server environment variables.
seedINT420–9223372036854776000Seeds both reference VAE encoding and target sampling.
nfe_stepsINT324–64Base AuK sampling steps. AuK-Flash requires 4.
cfg_strengthFLOAT2.00–5Base AuK classifier-free guidance. AuK-Flash requires 0.
sway_sampling_coefFLOAT-1.0-1–1Base AuK sway sampling coefficient. AuK-Flash requires the default -1 placeholder.
input_audiooptAUDIOReference audio for zero-shot TTS or source audio for editing. Leave unconnected for instruction TTS. Batch size must be 1; channels are averaged to mono.
llm_configoptAUK_LLM_CONFIG_DocConnect AuK OpenAI Settings_Doc or AuK Llama.cpp Adapter_Doc for Prompt Enhancer. Unconnected uses environment variables. Ignored when Prompt Enhancer is disabled.

Outputs (3)

NameTypeDescription
generated audioAUDIO—
model instructionSTRING—
Prompt Enhancer infoSTRING—