Nodes/ComfyUI-mnemic-nodes/✨ LLM Request
ComfyUI Node

✨ LLM Request

ChatGPT, Claude and Your Local Ollama Through a Single Wire

By MNeMoNiCuZ·Created 3 years ago·Updated about 12 hours ago· 105
✨ LLM Request
  • images
  • video
  • response
  • thinking
  • success
  • status
◄endpointOllama (this PC)►
◄model►
◄presetUse [system_message] and [user_input]►
◄system_message►
◄user_input►
◄temperature0.80►
◄reasoningdefault►
◄max_tokens0►
◄top_p1.00►
◄seed42►
◄stop►
◄json_modefalse►
◄unload_model_afterfalse►
◄context_length0►
◄free_comfy_vramfalse►
◄max_retries2►
◄raise_on_errortrue►
◄custom_endpoint►

If your model is a 2026 LLM-encoder checkpoint - Z-Image, Flux 2 Klein, Anima, Krea 2 - your prompt is an instruction read by a general-purpose model, not a bag of tokens. So have another LLM write that instruction: plain English in, the structured prompt out. This node sends a prompt, and optionally images, to just about any language model and returns the reply as a string you wire into a CLIP Text Encode.

Why you'd reach for it

Two crowds, two jobs. First, prompt enhancement - rough idea in, structured prompt out, batched and wired downstream. That stopped being a novelty: mentions of "prompt enhancer" went from 12 in 2023 to 253 in the first half of 2026, and it fits the architecture instead of fighting it, because an LLM writing an instruction for an LLM encoder is a translation between two things that speak the same language. Second, captioning: feed it images, and it's a dataset labeler - the job people run Qwen3-VL or Florence-2 for. If you already have a Claude or Gemini key, or an Ollama box with a VLM, this skips installing another captioner.

How it works

It's a thin front-end over a pile of server protocols, and the design decision that matters is this: nothing secret lives in your workflow. A workflow - and the image metadata made with it - stores only the endpoint's name and the model name. The URL, the key and any headers are looked up at run time from the pack's .env and its endpoint JSON, so keys never reach the browser, and anything a server echoes in an error is masked first.

Endpoints live in nodes/llm/DefaultEndpoints.json and your own UserEndpoints.json, and each speaks one of five protocols: openai (LM Studio, llama.cpp, vLLM and most cloud APIs), anthropic (Claude), ollama (the native API - that's what unlocks num_ctx and keep_alive), and claude_cli / codex_cli, which shell out to the Claude Code or Codex CLI and spend your subscription instead of an API key. It's self-correcting, too: when an endpoint rejects a parameter by name, the node drops it, retries, and reports the change in status.

The inputs that actually matter

Sixteen inputs is a lot; you touch four. endpoint picks the server. model is the name the endpoint knows it by - empty means the endpoint's default (a local server falls back to the first chat model it lists; for Ollama, one already in memory). user_input is the request, and system_message is the role-and-rules block, ignored while a preset is selected.

The rest are advanced; a few earn their keep. images feeds vision models - every image in the batch goes. temperature defaults to 0.8 and is dropped automatically where refused. reasoning maps one setting onto reasoning_effort, Ollama's think, or Claude's extended thinking, and json_mode asks for valid JSON with the code fences stripped. For local models, unload_model_after frees the Ollama model's VRAM the moment it replies, and free_comfy_vram unloads ComfyUI's models before the call - both exist because a local LLM and an image model fighting over one GPU is this workflow's most common regret.

Outputs: response is the reply with any <think> reasoning stripped out; thinking carries that reasoning separately; success is a boolean; status is the HTTP status or error text. Wire response into your prompt encoder. Turn off raise_on_error and a failed call returns an empty response and success = false instead of killing the queue - that boolean is your branch point.

Installing it

ComfyUI Manager is the easy path: search ComfyUI-mnemic-nodes and install. By hand:

cd ComfyUI/custom_nodes
git clone https://github.com/MNeMoNiCuZ/ComfyUI-mnemic-nodes

then restart ComfyUI. It's a grab-bag pack from MNeMoNiCuZ (mnemic2 on Reddit), so requirements.txt pulls a spread of utility deps - groq, transformers, tiktoken, python-dotenv, opencv-python, piexif. There are no model files to download, which is the appeal: cloud endpoints want a key, and Ollama and LM Studio on the same PC need no setup at all. Start ComfyUI once and it creates a .env from .env.example; fill in only what you use:

ANTHROPIC_API_KEY=sk-ant-...
OLLAMA_NETWORK_URL=http://192.168.1.50:11434

Edits apply on the next run, no restart. A server elsewhere on your network goes in UserEndpoints.json (press R to refresh), and the node's status panel shows where the endpoint runs - 🖥 this PC, 🏠 network, ☁ cloud - and whether its key is set.

Where people get burned

The node caches its reply while its inputs are unchanged and only runs when something downstream consumes its output - if an identical prompt seems stuck, bump the seed or change an input. A feature that reads as a bug exactly once. If your local model and your image model share a GPU, you'll see the swap-thrash: reach for unload_model_after and free_comfy_vram before you decide the node is slow.

Cloud endpoints refuse things - most commonly NSFW, which is why a chunk of the community runs a local model here instead of paying an API. The Custom Endpoint mode lets you type an address and key straight onto the node, stored in plain text in nodes/llm/CustomEndpoints.local.json - fine on your desktop, but don't leave a paid key there on a shared instance. And remember custom nodes are just Python with your user's permissions, so install from sources you trust.

Category⚡ MNeMiC Nodes

Inputs (20)

NameTypeDefaultDescription
endpointCOMBOOllama (this PC)Which server to talk to. Endpoints are defined in nodes/llm/*.json; keys and private addresses come from .env and are never saved in the workflow.
modelSTRINGModel name as the endpoint knows it. Empty uses the endpoint's default model, or on a local/network server the first model it lists (for Ollama, one already in memory). The Claude Code/Codex CLIs and some servers pick their own default. Use 🔍 Models on the node to browse.
presetCOMBOUse [system_message] and [user_input]A saved system prompt, shared with the Groq nodes. The first entry uses the system_message field instead.
system_messageSTRINGInstructions setting the model's role and rules. Ignored while a preset is selected.
user_inputSTRINGThe request itself: what you want the model to write, rewrite or describe. May be empty: the system message or preset is then sent on its own.
temperatureFLOAT0.800–2Randomness. Low is focused and repeatable, high is varied and creative. Dropped automatically for models that only allow their default.
reasoningCOMBOdefaultHow hard a reasoning model should think. 'default' sends nothing; 'none' turns thinking off where supported. On Claude this sets adaptive thinking's effort (a thinking budget on older models).
max_tokensINT00–262144Maximum length of the reply in tokens. 0 leaves it to the server (Claude, which needs a value, gets 16000).
top_pFLOAT1.000–1Nucleus sampling: only consider the most likely words adding up to this probability. 1.0 disables it and is not sent.
seedINT420–4294967295Sent to the endpoint for repeatable replies where supported. Randomized after each run; set control_after_generate to fixed to reuse the same reply.
stopSTRINGStop generating when this text appears. Separate several with |. Empty sends none.
json_modeBOOLEANfalseAsk for a valid JSON reply. Code fences around the reply are removed.
unload_model_afterBOOLEANfalseOllama only: unload the model from memory right after replying, so image generation gets the VRAM back.
context_lengthINT00–1048576Ollama only: context window in tokens (num_ctx). 0 uses the model's default.
free_comfy_vramBOOLEANfalseUnload ComfyUI's models from VRAM before the call. Useful when a local LLM shares the GPU with ComfyUI.
max_retriesINT20–10Extra attempts on connection errors, rate limits and server errors. 0 tries once.
raise_on_errorBOOLEANtrueStop the workflow with an error when the call fails. Off returns an empty response and success = false instead, for branching.
imagesoptIMAGEImages to send along with the prompt, for vision models. Every image in the batch is sent.
videooptVIDEOA Thousand Words only: a video to caption instead of, or alongside, images. Which models accept video depends on the server's own model list.
custom_endpointoptSTRINGOnly for 'Custom Endpoint - WARNING': an id set by the node's custom-endpoint panel. The address and key it points to are stored on this machine, never in the workflow. Empty for every other endpoint.

Outputs (4)

NameTypeDescription
responseSTRINGThe model's reply, with any <think> reasoning removed.
thinkingSTRINGThe model's reasoning, when the endpoint returns it. Empty otherwise.
successBOOLEANTrue when the call succeeded. Only false when raise_on_error is off.
statusSTRINGHTTP status or error message, e.g. '200 OK'. Never contains keys.