ComfyUI Node

LLM Connection

The one node that points ComfyUI at your local LLM

By VirusShell·Created 4 months ago·Updated 3 days ago· 1
LLM Connection
    • provider
    ◄urlhttp://localhost:1234►
    ◄host_mode▾►
    ◄model▾►
    ◄model_fallback►
    ◄timeout0►
    ◄manage_model_memorytrue►
    ◄ttl30►
    ◄context_length0►

    What it is, and the thing people get wrong first

    LLM Connection is a backend selector. That's the whole node. You point it at an LLM server running somewhere else - LM Studio, text-generation-webui, llama.cpp's llama-server, OpenAI, or anything that speaks /v1/chat/completions - and it hands the rest of your graph a provider object describing how to talk to that server. Add LLM Generate (Basic) downstream, wire provider into it, and you've got text generation inside ComfyUI.

    The misconception worth killing early: this does not replace your model's text encoder. A local LLM in a ComfyUI graph writes text - a prompt rewrite, a caption, dialogue for a video clip - and that text still goes into the checkpoint's own encoder (T5, Qwen3, whatever the base uses) exactly as if you'd typed it. LM Studio is not a Qwen encoder. It's a text generator you use to produce a better string.

    Also: there is no API key field. Keys live in config.yaml or environment variables and never touch workflow JSON - the right call, given that an LLM-shaped node is exactly the category that got weaponized once (ComfyUI_LLMVISION).

    How it works

    On queue, the node normalizes your URL - it strips duplicated /v1 suffixes, since chat and model-list code appends its own paths - then resolves a backend. If host_mode is pinned, that backend wins. On Auto (detect) it probes in parallel and fingerprints the server: Ollama's /api/version, llama.cpp's /health, LM Studio's /api/v1/models, Textgen's /v1/internal/model/info, falling back to openai if /v1/models answers and generic if nothing does.

    Out the other side comes a plain dict your Generate node consumes: backend, adapter: oai_compat, cleaned url, model, resolved timeout, plus an optional lifecycle block. All generation then goes through one adapter - POST {url}/v1/chat/completions. No provider SDKs; the only dependencies are pyyaml and requests.

    One thing to internalize: the VRAM knobs on this node are not doing anything to your GPU. They're instructions packaged into that dict and carried downstream - Textgen load/unload, LM Studio TTL - so the generation node can act on them.

    The fields you'll actually touch

    url - the base address, no /v1. LM Studio defaults to http://localhost:1234, Textgen http://localhost:5000.

    host_mode - six choices: Auto (detect), LM Studio, Textgen, llama.cpp, OpenAI / OAI-compat, Generic OAI. Leave it Auto unless detection guesses wrong. Pin llama.cpp or Generic OAI and the model field becomes free text.

    model - a dropdown populated from the server. Hit Refresh Models and pick one. If the catalog comes back empty, type the model id in by hand.

    model_fallback - a STRING socket, not a text box. Wire something into it and it overrides whatever the dropdown says. Handy when one workflow has to hit different model ids on different machines.

    Then, all optional and all mode-specific: manage_model_memory (Textgen only - load before generating, unload after the chain), ttl and context_length (LM Studio only), timeout (0 means use config, then fall back to 120), and ensure_load_on_select, which is a frontend-only convenience - the Python side literally ignores it.

    The single output is provider (type LLM_PROVIDER). It goes into LLM Generate (Basic) or (Advanced)'s provider input. That's the graph.

    Install

    Via Manager, search the registry listing @amvir/comfyui-llm-bikeshed. By hand:

    cd ComfyUI/custom_nodes
    git clone https://github.com/VirusShell/comfyui-llm-bikeshed.git
    cd comfyui-llm-bikeshed
    pip install -r requirements.txt
    

    Restart ComfyUI; the nodes land under the LLM Bikeshed category. No model downloads - this pack ships no weights, it talks to a server you already run. Then set keys if you need them:

    cp config.example.yaml config.yaml
    # edit config.yaml, then RESTART - there is no browser reload-config
    

    Env vars work too: LLM_BIKESHED_LM_STUDIO_API_KEY, LLM_BIKESHED_OPENAI_API_KEY, and so on. There's a starter workflow in example_workflows/connection_basic.json.

    Where people get burned

    Ollama is not supported natively. Auto-detect may label a server as Ollama, but this pack has no Ollama /api/chat path - Auto deliberately downgrades it to a generic OAI face. Use something that exposes /v1/chat/completions in front of it.

    Empty model = validation error. Queue with the dropdown still reading (refresh to load) and you get a friendly nudge to set a model id. Refresh first, or type an id.

    Knobs in the wrong mode are silently ignored. Setting ttl while you're pointed at Textgen does nothing, and manage_model_memory does nothing on LM Studio. The tooltips say so, but people don't read tooltips.

    Cancel doesn't clean up VRAM. Cancelling aborts the HTTP request in ComfyUI, and Textgen may stop host-side too, but there's no unload-on-cancel: your model stays loaded (Textgen) or follows the normal TTL (LM Studio). If you're juggling a diffusion model and an LLM on one card, that matters.

    API/headless mode is awkward. The model dropdown is populated by client-server routes that aren't ComfyUI API-mode compatible, so a headless queue needs the URL and model already baked into the workflow JSON.

    Old graphs, old nodes. The separate Provider and Lifecycle nodes still exist so old workflows keep loading. For anything new, use Connection - that's the pack author's own recommendation.

    CategoryLLM Bikeshed/providers

    Inputs (8)

    NameTypeDefaultDescription
    urlSTRINGhttp://localhost:1234—
    host_modeCOMBO6 options: Auto (detect), LM Studio, Textgen, llama.cpp, OpenAI / OAI-compat, Generic OAI
    modelCOMBO1 options: (refresh to load)
    model_fallbackoptSTRING—
    timeoutoptINT00–86400HTTP timeout seconds. 0 = use config for the effective backend, then oai_compat, then 120.
    manage_model_memoryoptBOOLEANtrueManage VRAM / Manage memory (default ON). Textgen: load before generate and unload after the chain. LM Studio: TTL and context length on load. OFF: pack does not load, unload, or set TTL. Hidden for OpenAI, generic, and llama.cpp (router /models/load + /models/unload are not verified here).
    ttloptINT30LM Studio only, while Manage VRAM is ON: seconds to keep the model loaded after each request. 0 = unload immediately. Ignored when Manage VRAM is OFF or the face is not LM Studio.
    context_lengthoptINT00–1048576LM Studio only, while Manage VRAM is ON: context window on explicit load. 0 = model default. Ignored when Manage VRAM is OFF or the face is not LM Studio.

    Outputs (1)

    NameTypeDescription
    providerLLM_PROVIDER—