✨ LLM Request
ChatGPT, Claude and Your Local Ollama Through a Single Wire
- images
- video
- response
- thinking
- success
- status
If your model is a 2026 LLM-encoder checkpoint - Z-Image, Flux 2 Klein, Anima, Krea 2 - your prompt is an instruction read by a general-purpose model, not a bag of tokens. So have another LLM write that instruction: plain English in, the structured prompt out. This node sends a prompt, and optionally images, to just about any language model and returns the reply as a string you wire into a CLIP Text Encode.
Why you'd reach for it
Two crowds, two jobs. First, prompt enhancement - rough idea in, structured prompt out, batched and wired downstream. That stopped being a novelty: mentions of "prompt enhancer" went from 12 in 2023 to 253 in the first half of 2026, and it fits the architecture instead of fighting it, because an LLM writing an instruction for an LLM encoder is a translation between two things that speak the same language. Second, captioning: feed it images, and it's a dataset labeler - the job people run Qwen3-VL or Florence-2 for. If you already have a Claude or Gemini key, or an Ollama box with a VLM, this skips installing another captioner.
How it works
It's a thin front-end over a pile of server protocols, and the design decision that matters is this: nothing secret lives in your workflow. A workflow - and the image metadata made with it - stores only the endpoint's name and the model name. The URL, the key and any headers are looked up at run time from the pack's .env and its endpoint JSON, so keys never reach the browser, and anything a server echoes in an error is masked first.
Endpoints live in nodes/llm/DefaultEndpoints.json and your own UserEndpoints.json, and each speaks one of five protocols: openai (LM Studio, llama.cpp, vLLM and most cloud APIs), anthropic (Claude), ollama (the native API - that's what unlocks num_ctx and keep_alive), and claude_cli / codex_cli, which shell out to the Claude Code or Codex CLI and spend your subscription instead of an API key. It's self-correcting, too: when an endpoint rejects a parameter by name, the node drops it, retries, and reports the change in status.
The inputs that actually matter
Sixteen inputs is a lot; you touch four. endpoint picks the server. model is the name the endpoint knows it by - empty means the endpoint's default (a local server falls back to the first chat model it lists; for Ollama, one already in memory). user_input is the request, and system_message is the role-and-rules block, ignored while a preset is selected.
The rest are advanced; a few earn their keep. images feeds vision models - every image in the batch goes. temperature defaults to 0.8 and is dropped automatically where refused. reasoning maps one setting onto reasoning_effort, Ollama's think, or Claude's extended thinking, and json_mode asks for valid JSON with the code fences stripped. For local models, unload_model_after frees the Ollama model's VRAM the moment it replies, and free_comfy_vram unloads ComfyUI's models before the call - both exist because a local LLM and an image model fighting over one GPU is this workflow's most common regret.
Outputs: response is the reply with any <think> reasoning stripped out; thinking carries that reasoning separately; success is a boolean; status is the HTTP status or error text. Wire response into your prompt encoder. Turn off raise_on_error and a failed call returns an empty response and success = false instead of killing the queue - that boolean is your branch point.
Installing it
ComfyUI Manager is the easy path: search ComfyUI-mnemic-nodes and install. By hand:
cd ComfyUI/custom_nodes
git clone https://github.com/MNeMoNiCuZ/ComfyUI-mnemic-nodes
then restart ComfyUI. It's a grab-bag pack from MNeMoNiCuZ (mnemic2 on Reddit), so requirements.txt pulls a spread of utility deps - groq, transformers, tiktoken, python-dotenv, opencv-python, piexif. There are no model files to download, which is the appeal: cloud endpoints want a key, and Ollama and LM Studio on the same PC need no setup at all. Start ComfyUI once and it creates a .env from .env.example; fill in only what you use:
ANTHROPIC_API_KEY=sk-ant-...
OLLAMA_NETWORK_URL=http://192.168.1.50:11434
Edits apply on the next run, no restart. A server elsewhere on your network goes in UserEndpoints.json (press R to refresh), and the node's status panel shows where the endpoint runs - 🖥 this PC, 🏠 network, ☁ cloud - and whether its key is set.
Where people get burned
The node caches its reply while its inputs are unchanged and only runs when something downstream consumes its output - if an identical prompt seems stuck, bump the seed or change an input. A feature that reads as a bug exactly once. If your local model and your image model share a GPU, you'll see the swap-thrash: reach for unload_model_after and free_comfy_vram before you decide the node is slow.
Cloud endpoints refuse things - most commonly NSFW, which is why a chunk of the community runs a local model here instead of paying an API. The Custom Endpoint mode lets you type an address and key straight onto the node, stored in plain text in nodes/llm/CustomEndpoints.local.json - fine on your desktop, but don't leave a paid key there on a shared instance. And remember custom nodes are just Python with your user's permissions, so install from sources you trust.
Inputs (20)
| Name | Type | Default | Description |
|---|---|---|---|
| endpoint | COMBO | Ollama (this PC) | Which server to talk to. Endpoints are defined in nodes/llm/*.json; keys and private addresses come from .env and are never saved in the workflow. |
| model | STRING | Model name as the endpoint knows it. Empty uses the endpoint's default model, or on a local/network server the first model it lists (for Ollama, one already in memory). The Claude Code/Codex CLIs and some servers pick their own default. Use 🔍 Models on the node to browse. | |
| preset | COMBO | Use [system_message] and [user_input] | A saved system prompt, shared with the Groq nodes. The first entry uses the system_message field instead. |
| system_message | STRING | Instructions setting the model's role and rules. Ignored while a preset is selected. | |
| user_input | STRING | The request itself: what you want the model to write, rewrite or describe. May be empty: the system message or preset is then sent on its own. | |
| temperature | FLOAT | 0.800–2 | Randomness. Low is focused and repeatable, high is varied and creative. Dropped automatically for models that only allow their default. |
| reasoning | COMBO | default | How hard a reasoning model should think. 'default' sends nothing; 'none' turns thinking off where supported. On Claude this sets adaptive thinking's effort (a thinking budget on older models). |
| max_tokens | INT | 00–262144 | Maximum length of the reply in tokens. 0 leaves it to the server (Claude, which needs a value, gets 16000). |
| top_p | FLOAT | 1.000–1 | Nucleus sampling: only consider the most likely words adding up to this probability. 1.0 disables it and is not sent. |
| seed | INT | 420–4294967295 | Sent to the endpoint for repeatable replies where supported. Randomized after each run; set control_after_generate to fixed to reuse the same reply. |
| stop | STRING | Stop generating when this text appears. Separate several with |. Empty sends none. | |
| json_mode | BOOLEAN | false | Ask for a valid JSON reply. Code fences around the reply are removed. |
| unload_model_after | BOOLEAN | false | Ollama only: unload the model from memory right after replying, so image generation gets the VRAM back. |
| context_length | INT | 00–1048576 | Ollama only: context window in tokens (num_ctx). 0 uses the model's default. |
| free_comfy_vram | BOOLEAN | false | Unload ComfyUI's models from VRAM before the call. Useful when a local LLM shares the GPU with ComfyUI. |
| max_retries | INT | 20–10 | Extra attempts on connection errors, rate limits and server errors. 0 tries once. |
| raise_on_error | BOOLEAN | true | Stop the workflow with an error when the call fails. Off returns an empty response and success = false instead, for branching. |
| imagesopt | IMAGE | Images to send along with the prompt, for vision models. Every image in the batch is sent. | |
| videoopt | VIDEO | A Thousand Words only: a video to caption instead of, or alongside, images. Which models accept video depends on the server's own model list. | |
| custom_endpointopt | STRING | Only for 'Custom Endpoint - WARNING': an id set by the node's custom-endpoint panel. The address and key it points to are stored on this machine, never in the workflow. Empty for every other endpoint. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| response | STRING | The model's reply, with any <think> reasoning removed. |
| thinking | STRING | The model's reasoning, when the endpoint returns it. Empty otherwise. |
| success | BOOLEAN | True when the call succeeded. Only false when raise_on_error is off. |
| status | STRING | HTTP status or error message, e.g. '200 OK'. Never contains keys. |