llama.cpp ADV++ Prompt
The ADV++ kitchen sink
- trigger
- token_ban
- image_1
- image_2
- image_3
- image_4
- image_5
- image_6
- image_7
- image_8
- image_9
- image_10
- structured_output
- connection
- response
- thinking
- success
ADV Prompt does text-plus-images. ADV++ takes that and stacks on the three things that turn a chatty LLM into a reliable pipeline component: bundled prompt templates, token bans, and structured-output constraints. If you've ever gotten a caption back that starts with "Here is a detailed description of your image," this node is the cure.
It's the node you reach for when the LLM isn't just generating text - it's doing a job with strict output requirements. Caption → image prompt. Rough idea → structured prompt. Raw text → valid JSON. Those jobs are exactly where a plain chat model fails, because it hands you its conversational habits as literal output.
How it works
ADV++ is ADV Prompt's generation engine with three extra plug-in points, each coming from a sibling node or a bundled file:
- Templates - the
templatedropdown applies a bundled prompt template before generation. The pack shipsImage2Prompt(rewrite an image into a single final image-generation prompt) andPrompt Enhancer(turn rough ideas into detailed prompts), defined inweb/templates.json. Each template injects asystem_promptand an optional default user prompt. Templates are applied in Python, so API-format and headless workflows behave the same as the UI. - Token bans - a
token_baninput takes a list from a llama.cpp Token Ban node. The bans are sent as llama.cpp text-form logit-bias entries, so generation is steered away from specific tokens the same way the LTX-style hard-stop nodes block role delimiters. - Structured output - a
structured_outputinput accepts a constraint built by llama.cpp Structured Output (JSON Schema, JSON object, or GBNF grammar), which forces decoding into the shape you asked for.
The inputs that matter
Most of the sampling controls match ADV Prompt, so I'll skip those and hit the distinctive ones:
- template -
Empty,Image2Prompt, orPrompt Enhancer. Restart ComfyUI after editingweb/templates.jsonto add your own. - prompt - the user prompt. With Image2Prompt this is where an image description or
output:marker goes. - image_amount / image_1…image_10 / include_image_batch - the full vision stack, same as ADV Prompt.
- token_ban + enable_token_ban - plug in a Token Ban node and toggle it.
- structured_output - plug in a Structured Output node.
enableon that node is what bypasses the constraint when you're experimenting. - stop_sequences - still useful even with structured output; belt and braces.
Outputs: response (the final generated or structured text), thinking, and success.
Why this matters
The KB's LLM-in-ComfyUI essay is blunt about the failure mode: a chat LLM doesn't emit a clean prompt by default - it emits role delimiters, preamble, markdown scaffolding. ADV++ gives you the three-layer fix in one node: a template that tells it the exact job, token bans that physically stop it from writing certain tokens, and structured output that constrains decoding itself. That combination is the difference between an enhancer that makes your prompts worse and one that earns its place in the graph.
Common issues
- Template output looks wrong - the bundled templates are opinionated (Image2Prompt is explicitly NSFW-tolerant and refuses to omit visible anatomy). Edit
web/templates.jsonand restart if you want your own rules. - Structured output and token bans fighting - a ban that contradicts the grammar can stall generation. Disable one while you debug.
- Model adds text outside the JSON - make sure
strictis on in the Structured Output node andsystem_promptis empty or minimal; the template's system prompt can override your formatting intent.
Inputs (38)
| Name | Type | Default | Description |
|---|---|---|---|
| template | COMBO | Empty | Apply a bundled prompt template before generation. |
| prompt | STRING | The user prompt to send to the LLM | |
| image_amount | INT | 20–10 | Number of image input slots to show |
| modelopt | COMBO | (use running model) | Model for router mode, or the running direct model. |
| server_urlopt | STRING | Leave empty to use the server owned by this node pack. Attached endpoints are never implicitly stopped. | |
| system_promptopt | STRING | System prompt that defines model behavior. | |
| enable_thinkingopt | BOOLEAN | true | Request thinking/reasoning from compatible models. |
| max_tokensopt | INT | 20481–131072 | Maximum number of tokens to generate. |
| temperatureopt | FLOAT | 0.700–2 | Sampling randomness. Lower values are more deterministic. |
| top_popt | FLOAT | 0.900–1 | Keep tokens within this cumulative probability mass. |
| top_kopt | INT | 400–200 | Sample from the top K tokens. 0 disables top-k filtering. |
| min_popt | FLOAT | 0.050–1 | Discard tokens below this probability relative to the best token. |
| repeat_penaltyopt | FLOAT | 1.101–2 | Penalize recently repeated tokens. 1.0 disables the penalty. |
| presence_penaltyopt | FLOAT | 0.0-2–2 | Penalize tokens that have appeared at least once. |
| frequency_penaltyopt | FLOAT | 0.0-2–2 | Penalize tokens in proportion to how often they appeared. |
| seedopt | INT | 00–2147483647 | Random seed |
| keep_contextopt | BOOLEAN | false | Reuse a matching prompt-prefix KV cache. This is not chat history. |
| enable_chainingopt | BOOLEAN | false | Compatibility toggle. A connected trigger already controls ordering. |
| triggeropt | * | Optional dependency input used to sequence execution. | |
| token_banopt | LOGIT_BIAS | Token ban list from a llama.cpp Token Ban node. | |
| enable_token_banopt | BOOLEAN | true | Enable or disable the connected token ban list. |
| stop_sequencesopt | STRING | Stop sequences. JSON arrays preserve commas and whitespace. | |
| api_key_envopt | STRING | LLAMACPP_API_KEY | Environment variable containing the API key. The secret is not serialized. |
| verify_tlsopt | BOOLEAN | true | Verify HTTPS certificates. |
| request_timeoutopt | INT | 3001–86400 | Overall generation deadline in seconds. |
| include_image_batchopt | BOOLEAN | false | Send every image in each connected IMAGE batch. |
| image_1opt | IMAGE | Optional image 1. Visibility follows image_amount. | |
| image_2opt | IMAGE | Optional image 2. Visibility follows image_amount. | |
| image_3opt | IMAGE | Optional image 3. Visibility follows image_amount. | |
| image_4opt | IMAGE | Optional image 4. Visibility follows image_amount. | |
| image_5opt | IMAGE | Optional image 5. Visibility follows image_amount. | |
| image_6opt | IMAGE | Optional image 6. Visibility follows image_amount. | |
| image_7opt | IMAGE | Optional image 7. Visibility follows image_amount. | |
| image_8opt | IMAGE | Optional image 8. Visibility follows image_amount. | |
| image_9opt | IMAGE | Optional image 9. Visibility follows image_amount. | |
| image_10opt | IMAGE | Optional image 10. Visibility follows image_amount. | |
| structured_outputopt | STRUCTURED_OUTPUT | JSON schema, JSON object, or GBNF constraint. | |
| connectionopt | LLAMACPP_CONNECTION | Optional reusable local or remote connection profile. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| response | STRING | Generated multimodal or structured response text. |
| thinking | STRING | Reasoning content reported separately by compatible models. |
| success | BOOLEAN | Whether generation completed successfully. |