ComfyUI Node
llama.cpp Basic Prompt
Runs local text generation through llama-server with sampling, reasoning, and workflow-chaining controls.
llama.cpp Basic Prompt
- trigger
- connection
- response
- thinking
- success
◄prompt►
◄model(use running model)►
◄server_url►
◄system_prompt►
◄enable_thinkingtrue►
◄max_tokens2048►
◄temperature0.70►
◄top_p0.90►
◄top_k40►
◄min_p0.05►
◄repeat_penalty1.10►
◄seed0►
◄keep_contextfalse►
◄enable_chainingfalse►
◄presence_penalty0.0►
◄frequency_penalty0.0►
◄stop_sequences►
◄api_key_envLLAMACPP_API_KEY►
◄verify_tlstrue►
◄request_timeout300►
CategoryLlamaCpp
Inputs (22)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | The user prompt to send to the LLM | |
| modelopt | COMBO | (use running model) | Model for router mode, or the running direct model. |
| server_urlopt | STRING | Leave empty to use the server owned by this node pack. Attached endpoints are never implicitly stopped. | |
| system_promptopt | STRING | System prompt that defines model behavior. | |
| enable_thinkingopt | BOOLEAN | true | Request thinking/reasoning from compatible models. |
| max_tokensopt | INT | 20481–131072 | Maximum number of tokens to generate. |
| temperatureopt | FLOAT | 0.700–2 | Sampling randomness. Lower values are more deterministic. |
| top_popt | FLOAT | 0.900–1 | Keep tokens within this cumulative probability mass. |
| top_kopt | INT | 400–200 | Sample from the top K tokens. 0 disables top-k filtering. |
| min_popt | FLOAT | 0.050–1 | Discard tokens below this probability relative to the best token. |
| repeat_penaltyopt | FLOAT | 1.101–2 | Penalize recently repeated tokens. 1.0 disables the penalty. |
| seedopt | INT | 00–2147483647 | Random seed |
| keep_contextopt | BOOLEAN | false | Reuse a matching prompt-prefix KV cache. This is not chat history. |
| enable_chainingopt | BOOLEAN | false | Compatibility toggle. A connected trigger already controls ordering. |
| triggeropt | * | Optional dependency input used to sequence execution. | |
| presence_penaltyopt | FLOAT | 0.0-2–2 | Penalize tokens that have appeared at least once. |
| frequency_penaltyopt | FLOAT | 0.0-2–2 | Penalize tokens in proportion to how often they appeared. |
| stop_sequencesopt | STRING | Stop sequences. JSON arrays preserve commas and whitespace. | |
| api_key_envopt | STRING | LLAMACPP_API_KEY | Environment variable containing the API key. The secret is not serialized. |
| verify_tlsopt | BOOLEAN | true | Verify HTTPS certificates. |
| request_timeoutopt | INT | 3001–86400 | Overall generation deadline in seconds. |
| connectionopt | LLAMACPP_CONNECTION | Optional reusable local or remote connection profile. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| response | STRING | Generated response text. |
| thinking | STRING | Reasoning content reported separately by compatible models. |
| success | BOOLEAN | Whether generation completed successfully. |