AuK Models Loader (Low VRAM)_Doc
Getting AuK to fit on a normal GPU
- engine
This is the boring node at the front of every AuK workflow, and it's the one that decides whether you get a run or an out-of-memory crash. Tencent's AuK wants a DiT checkpoint and a 3B Qwen2.5-Omni encoder resident at once, which is a lot to ask of a card that's already holding a video model. The loader's entire job is to let you pick which of those two gets the GPU and which gets parked on CPU.
What it actually loads
Two things, and both are chosen from dropdowns that scan your models folder:
model_name- an AuK or AuK-Flash checkpoint. Discovery isn't a file glob; it's a directory scan for a.safetensorsthat has aconfig.yamland avae.safetensorssitting beside it, insideComfyUI/models/aukorComfyUI/models/diffusion_models/auk. Miss any one of the three files and the checkpoint simply won't appear in the list.qwen_name- a complete Qwen2.5-Omni directory, found by readingconfig.jsonfiles underComfyUI/models/text_encoders(andmodels/LLM) and keeping the ones whosemodel_typestarts withqwen2_5_omni. It must be the full original model: index, all three shards, tokenizer and processor configs. A GGUF quant of it won't be listed, which is the point - GGUF in this pack is only for the optional llama.cpp prompt-enhancer path.
device lists your actual CUDA devices, and the node refuses to run without one. dtype is bf16 or fp16; if your card doesn't report bf16 support you get a clear error telling you to pick fp16, which is friendlier than most packs manage.
memory_mode is the whole reason you're here
low_vram (the default) keeps Qwen and the VAE on CPU and puts a half-precision DiT on the GPU, offloading layers in as needed. balanced bounces Qwen and the DiT on and off the GPU and leaves only the VAE on CPU.
The author's own validation report tells you which to pick, which is rare and welcome: on an RTX 5070 Ti capped at a 7 GiB PyTorch allocation, low_vram finished all three test generations at roughly 2.95 GiB allocated / 3.03 GiB reserved, with the whole card peaking around 6.3 GiB (including whatever else was already resident) and process RSS peaking near 27 GiB. Same tests under balanced OOM'd at a 7 GiB cap and only passed with a 12 GiB cap, at ~7.7 GiB allocated - and it wasn't faster. So: 8 GiB card, or any card you also want to do other things with, use low_vram. balanced is for people with VRAM to spare and no patience.
sequential_cfg (on by default) evaluates the two classifier-free-guidance branches one after the other instead of together. Less activation memory, more wall-clock. You're already memory-starved if you're reading this page; leave it on.
Wiring
One output: engine (AUK_ENGINE_Doc), straight into AuK Generate / Edit_Doc's engine socket. It's a custom type, so only the Generate node will accept it - no adapters, no reuse elsewhere.
Worth knowing: the node caches the engine per unique combination of checkpoint, Qwen path, memory mode, precision, device and sequential-CFG setting, and calls ComfyUI's unload_all_models() plus a soft cache empty before building it. That eviction is deliberate - it's how AuK gets the room it needs - but it means a graph mixing AuK with a video model will reload everything each time you alternate. Don't assume the loader is broken when your checkpoint reloads.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/DocWorkBox/ComfyUI-AuK_Doc
cd ComfyUI-AuK_Doc
python -m pip install -r requirements.txt
Restart ComfyUI afterwards, using the same Python environment ComfyUI runs in. Or just search ComfyUI-AuK_Doc in ComfyUI Manager. The dependency list is not small (transformers>=4.52,<5, qwen-omni-utils, funasr, modelscope, silero-vad, torchdiffeq, openai, soundfile), and it intentionally leaves Torch alone because ComfyUI supplies it.
Models go here:
ComfyUI/models/
├─ auk/
│ ├─ AuK/ auk_base.safetensors + vae.safetensors + config.yaml
│ └─ AuK-Flash/ auk_flash.safetensors + vae.safetensors + config.yaml
└─ text_encoders/
└─ Qwen2.5-Omni-3B/ full directory: index, 3 shards, tokenizer, processor
Base and Flash share the Qwen encoder, so downloading both checkpoints costs you one extra model, not two. Weights are on ModelScope under Tencent-Hunyuan/AuK and Tencent-Hunyuan/AuK-Flash; the pack does no downloading at all - it only reads what's on disk.
When the dropdowns are empty
- No AuK models found - the scan needs all three files in one folder.
vae.safetensorslying next toauk_base.safetensorswithoutconfig.yamlis invisible. - No Qwen2.5-Omni models found - either
config.jsonreports a differentmodel_type, or you only copied the shards. It reads the config, not the weights. - "The selected configuration is not AuK or AuK-Flash" - you pointed it at some other
config.yaml. The loader checks the model name inside. - Model placed, still not listed - restart (or refresh) so
folder_pathspicks up the new directory.
Inputs (6)
| Name | Type | Default | Description |
|---|---|---|---|
| model_name | COMBO | 1 options: No AuK models found | |
| qwen_name | COMBO | 1 options: No Qwen2.5-Omni models found | |
| memory_mode | COMBO | low_vram | low_vram: Qwen + VAE on CPU, half-precision DiT on GPU. balanced: Qwen and DiT alternate on GPU, VAE on CPU; requires more VRAM and is not recommended for 8 GiB. |
| dtype | COMBO | bf16 | 2 options: bf16, fp16 |
| device | COMBO | 1 options: CUDA unavailable | |
| sequential_cfg | BOOLEAN | true | Evaluate CFG branches sequentially, reducing activation memory while taking more time. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| engine | AUK_ENGINE_Doc | — |