Nodes/ComfyUI-AuK_Doc/AuK Models Loader (Low VRAM)_Doc
ComfyUI Node

AuK Models Loader (Low VRAM)_Doc

Getting AuK to fit on a normal GPU

By DocWorkBox·Created 23 days ago·Updated 23 days ago· 4
AuK Models Loader (Low VRAM)_Doc
    • engine
    ◄model_name▾►
    ◄qwen_name▾►
    ◄memory_modelow_vram►
    ◄dtypebf16►
    ◄device▾►
    ◄sequential_cfgtrue►

    This is the boring node at the front of every AuK workflow, and it's the one that decides whether you get a run or an out-of-memory crash. Tencent's AuK wants a DiT checkpoint and a 3B Qwen2.5-Omni encoder resident at once, which is a lot to ask of a card that's already holding a video model. The loader's entire job is to let you pick which of those two gets the GPU and which gets parked on CPU.

    What it actually loads

    Two things, and both are chosen from dropdowns that scan your models folder:

    • model_name - an AuK or AuK-Flash checkpoint. Discovery isn't a file glob; it's a directory scan for a .safetensors that has a config.yaml and a vae.safetensors sitting beside it, inside ComfyUI/models/auk or ComfyUI/models/diffusion_models/auk. Miss any one of the three files and the checkpoint simply won't appear in the list.
    • qwen_name - a complete Qwen2.5-Omni directory, found by reading config.json files under ComfyUI/models/text_encoders (and models/LLM) and keeping the ones whose model_type starts with qwen2_5_omni. It must be the full original model: index, all three shards, tokenizer and processor configs. A GGUF quant of it won't be listed, which is the point - GGUF in this pack is only for the optional llama.cpp prompt-enhancer path.

    device lists your actual CUDA devices, and the node refuses to run without one. dtype is bf16 or fp16; if your card doesn't report bf16 support you get a clear error telling you to pick fp16, which is friendlier than most packs manage.

    memory_mode is the whole reason you're here

    low_vram (the default) keeps Qwen and the VAE on CPU and puts a half-precision DiT on the GPU, offloading layers in as needed. balanced bounces Qwen and the DiT on and off the GPU and leaves only the VAE on CPU.

    The author's own validation report tells you which to pick, which is rare and welcome: on an RTX 5070 Ti capped at a 7 GiB PyTorch allocation, low_vram finished all three test generations at roughly 2.95 GiB allocated / 3.03 GiB reserved, with the whole card peaking around 6.3 GiB (including whatever else was already resident) and process RSS peaking near 27 GiB. Same tests under balanced OOM'd at a 7 GiB cap and only passed with a 12 GiB cap, at ~7.7 GiB allocated - and it wasn't faster. So: 8 GiB card, or any card you also want to do other things with, use low_vram. balanced is for people with VRAM to spare and no patience.

    sequential_cfg (on by default) evaluates the two classifier-free-guidance branches one after the other instead of together. Less activation memory, more wall-clock. You're already memory-starved if you're reading this page; leave it on.

    Wiring

    One output: engine (AUK_ENGINE_Doc), straight into AuK Generate / Edit_Doc's engine socket. It's a custom type, so only the Generate node will accept it - no adapters, no reuse elsewhere.

    Worth knowing: the node caches the engine per unique combination of checkpoint, Qwen path, memory mode, precision, device and sequential-CFG setting, and calls ComfyUI's unload_all_models() plus a soft cache empty before building it. That eviction is deliberate - it's how AuK gets the room it needs - but it means a graph mixing AuK with a video model will reload everything each time you alternate. Don't assume the loader is broken when your checkpoint reloads.

    Install

    cd ComfyUI/custom_nodes
    git clone https://github.com/DocWorkBox/ComfyUI-AuK_Doc
    cd ComfyUI-AuK_Doc
    python -m pip install -r requirements.txt
    

    Restart ComfyUI afterwards, using the same Python environment ComfyUI runs in. Or just search ComfyUI-AuK_Doc in ComfyUI Manager. The dependency list is not small (transformers>=4.52,<5, qwen-omni-utils, funasr, modelscope, silero-vad, torchdiffeq, openai, soundfile), and it intentionally leaves Torch alone because ComfyUI supplies it.

    Models go here:

    ComfyUI/models/
    ├─ auk/
    │  ├─ AuK/          auk_base.safetensors + vae.safetensors + config.yaml
    │  └─ AuK-Flash/    auk_flash.safetensors + vae.safetensors + config.yaml
    └─ text_encoders/
       └─ Qwen2.5-Omni-3B/   full directory: index, 3 shards, tokenizer, processor
    

    Base and Flash share the Qwen encoder, so downloading both checkpoints costs you one extra model, not two. Weights are on ModelScope under Tencent-Hunyuan/AuK and Tencent-Hunyuan/AuK-Flash; the pack does no downloading at all - it only reads what's on disk.

    When the dropdowns are empty

    • No AuK models found - the scan needs all three files in one folder. vae.safetensors lying next to auk_base.safetensors without config.yaml is invisible.
    • No Qwen2.5-Omni models found - either config.json reports a different model_type, or you only copied the shards. It reads the config, not the weights.
    • "The selected configuration is not AuK or AuK-Flash" - you pointed it at some other config.yaml. The loader checks the model name inside.
    • Model placed, still not listed - restart (or refresh) so folder_paths picks up the new directory.
    CategoryAuK_Doc

    Inputs (6)

    NameTypeDefaultDescription
    model_nameCOMBO1 options: No AuK models found
    qwen_nameCOMBO1 options: No Qwen2.5-Omni models found
    memory_modeCOMBOlow_vramlow_vram: Qwen + VAE on CPU, half-precision DiT on GPU. balanced: Qwen and DiT alternate on GPU, VAE on CPU; requires more VRAM and is not recommended for 8 GiB.
    dtypeCOMBObf162 options: bf16, fp16
    deviceCOMBO1 options: CUDA unavailable
    sequential_cfgBOOLEANtrueEvaluate CFG branches sequentially, reducing activation memory while taking more time.

    Outputs (1)

    NameTypeDescription
    engineAUK_ENGINE_Doc—