AuK Llama.cpp Adapter_Doc
Local prompt enhancement, no key and no server
- llama_model
- llm_config
AuK's Prompt Enhancer wants a chat model to rewrite your request into the official instruction template. Most people wire up an OpenAI-compatible endpoint and call it done. This node is the other path: point it at a GGUF file you already have on disk, loaded by someone else's loader, and the enhancement happens entirely in-graph - no API key, no base_url, no second process listening on a port.
It's a small node with one socket, and it's the fiddliest install in the pack. Both of those things are worth knowing before you start.
How it works
The trick is that it doesn't load a model. llama-cpp_vllm's Llama-cpp Model Loader owns the model and keeps it in a module-level store; this adapter reaches into that store, hands it a chat-completions request, and wraps the reply in the same response object the OpenAI client would have returned. From the Prompt Enhancer's point of view, nothing is different - it thinks it's talking to a server. That's why the whole multi-stage enhancement flow still runs: classify the request, rewrite the instruction, estimate duration.
Wire it as Llama-cpp Model Loader → AuK Llama.cpp Adapter_Doc → AuK Generate / Edit_Doc.llm_config, then turn on Generate's Enable Prompt Enhancer. It's ignored when the enhancer is off, same as the OpenAI settings node.
The adapter uses the loader's shared instance, so it only reloads when the config it was handed doesn't match what's resident. And it holds a lock while it calls, so two runs can't fight over the same model.
The inputs
llama_model (LLAMACPPMODEL) is the only required wire - output of the Llama-cpp loader. Give it anything that isn't a loader model dict and it refuses rather than failing weirdly later.
temperature defaults to 0, which is the right default: this is a structured rewrite job, not a writing job, and a warm model here means rewritten instructions that wander off the template. Both it and max_tokens are range-checked (0–2, and 1–131072).
max_tokens at 4096 is headroom, not a target. A thinking-mode GGUF will treat it as a suggestion and spend the whole budget deliberating, at which point the enhancer gets non-JSON back and errors. Pick an instruct model, or a non-thinking mode.
unload_after_enhance (on by default) is the load-bearing one. After enhancement - successfully or not - it releases the shared llama.cpp model, and the loader reloads it when the next run needs it. That's what makes this viable on one GPU (llm-in-comfyui.md § 3 calls it out as the pattern good nodes converged on): you're budgeting VRAM for a language model and AuK at once, and something has to get out of the way. Leave it on unless you're running a big card and re-enhancement latency is hurting.
Output is a single llm_config, and it carries the model config with it - so nothing needs to be re-selected downstream.
Install: two extra things, and neither is in requirements.txt
cd ComfyUI/custom_nodes
git clone https://github.com/DocWorkBox/ComfyUI-AuK_Doc
cd ComfyUI-AuK_Doc
python -m pip install -r requirements.txt
That gets you the pack. Then, separately:
- Install the llama-cpp_vllm plugin (
lihaoyun6/ComfyUI-llama-cpp_vllm) - the pack's README is explicit that it needs a loader with theLLAMACPPMODELoutput type, and that other loaders with the same name aren't guaranteed to work. The adapter resolves the loader's storage by node id, so a lookalike will fail. - Make sure
llama-cpp-pythonactually imports. The pack deliberately does not touch it - requirements.txt says it keeps whatever CUDA build you have. This is the step that eats the afternoon.
On the second point, the community's experience is worth repeating because it saves you the search: llama-cpp-python doesn't publish wheels for the Python ComfyUI ships, so you grab a prebuilt wheel matching your Python and CUDA tags and pip it into the interpreter directly - not into custom_nodes. In this r/comfyui thread, the fix that worked was a wheel from the third-party builds at JamePeng/llama-cpp-python releases, with the tags read as cp313 = Python 3.13 and cu130 = CUDA 13.0. If the loader node shows as missing after a restart, this is why, and it's not an AuK problem.
Put your GGUF instruction model in ComfyUI/models/LLM and select it in the loader.
Troubleshooting
- "需要安装并启用 llama-cpp_vllm 的 Llama-cpp Model Loader" - the error text is in Chinese, but it means the adapter couldn't find a compatible loader on the node mappings. Plugin missing, disabled after a failed import, or the wrong loader. Fix the loader, not the adapter.
- "本地 llama.cpp 调用失败" - the model ran but errored. Usually context. The README's note applies: the loader's
vram_limitis approximate, and a0there does not mean CPU-only. Raisen_ctxif long instructions get truncated. - Non-JSON output / enhancement fails - the enhancer needs parseable JSON from the model. Tiny GGUFs and heavy reasoning models are the two usual culprits; a mid-size instruct model is the sweet spot.
- Your input audio's transcription still goes to Tencent - this node replaces the text LLM only. Any ASR of reference or source audio runs the AuK path unchanged (Tencent Cloud recording ASR with credentials, local SenseVoiceSmall otherwise). A local GGUF doesn't make the audio side offline.
- Re-loading every run is normal. With
unload_after_enhanceon, the loader re-reads the GGUF on the next enhancement. That's the trade for not holding two models in VRAM.
Inputs (4)
| Name | Type | Default | Description |
|---|---|---|---|
| llama_model | LLAMACPPMODEL | — | |
| temperature | FLOAT | 0.000–2 | — |
| max_tokens | INT | 40961–131072 | — |
| unload_after_enhance | BOOLEAN | true | Release the shared llama.cpp model after enhancement, including on failure. It reloads when next needed. |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| llm_config | AUK_LLM_CONFIG_Doc | — |