Nodes/ComfyUI-OpenRouter-API/Shubz - OpenRouter API
ComfyUI Node

Shubz - OpenRouter API

Media in, text out, and one default you should change first

By theshubzworld·Created a day ago·Updated a day ago· 1
Shubz - OpenRouter API
  • image
  • image_2
  • image_3
  • video
  • video_2
  • video_3
  • audio
  • audio_2
  • audio_3
  • text
  • info
  • credits
◄model— choose a compatible OpenRouter model —►
◄reasoning_effortauto►
◄timeout_seconds120►
◄temperature1.00►
◄max_tokens8192►
◄response_formattext►
◄zdrfalse►
◄regeneratetrue►
◄system_presetRef2VA►
◄system_prompt►
◄user_promptDescribe the supplied context.►
◄api_key►

This node doesn't generate anything

Look at the outputs before anything else: text, info, credits. No IMAGE, no VIDEO, no AUDIO - the request it sends hardcodes modalities: ["text"]. It's an OpenRouter chat client wearing a node costume: images, video or audio in, a STRING out that you feed to a text encoder or a video-API node.

That's the LLM-in-your-graph pattern - captioning a reference image, describing a clip before the video model sees it, reshaping a rough idea into something a checkpoint's encoder understands. Same job as a local prompt enhancer or captioner, except the thinking happens on OpenRouter's side and costs per call instead of VRAM. Why this over Comfy's official API nodes, which route through api.comfy.org and spend Comfy credits? Because this one takes your own key, no login.

The default will surprise you, so change it first

system_preset ships on Ref2VA. That's not a cute label - it injects a long MiniMax/Hailuo video prompt-writing guide as the system message, the full-reference dialect with subject_definitions, retention_analysis, detailed_description and overall_soundscape sections. Six options: Ref2VA, I2VA, RefAud2VA, T2VA, FL2VA, None. The "VA" ones target video models that generate picture and audio together - the MiniMax H3 lineage, where native stereo audio comes out of the same pass.

So connect a photo, leave user_prompt at its default Describe the supplied context., hit queue, and you don't get a caption. You get a structured audiovisual timeline for a video that doesn't exist - nice output, wrong job. Captioning, tagging, translating, rewriting: set system_preset to None. Anything you type in system_prompt gets appended under the preset after an [Additional Instructions] header, so it refines the preset rather than replacing it.

How it actually works

The request is an ordinary chat completion: an optional system message, then a user message made of your user_prompt plus each attachment as a base64 data URL. Before that POST leaves, every attachment gets compressed locally against a cap measured before base64 - images to WebP under 1,000,000 bytes, video to MP4/H.264/AAC under 10,000,000, audio to MP3 under 1,000,000 - with quality dropping before dimensions and an already-small compatible MP4 passed through untouched. The paid request doesn't start until everything is under cap, so an attachment that can't fit fails locally instead of being rejected upstream.

The model list is live: OpenRouter's metadata, filtered to text-output models, cached for an hour with a disk fallback marked stale. The front end intersects that with what you wired - text plus every connected kind among image/video/audio - which is why the dropdown shifts the moment you plug in a video. If your pick stops being compatible it resets to choose a compatible OpenRouter model rather than silently moving you onto a different paid model.

Parameters are gated on what the model advertises: temperature goes out only if supported, and an explicit reasoning_effort on a non-reasoning model or json_object without structured-output support fails before submission. timeout_seconds is one deadline covering catalog refresh, preprocessing and the chat call.

Inputs you'll actually touch

  • user_prompt - the message that travels with your media. Change the default.
  • system_preset - the video-prompt guides, or None (see above).
  • model and reasoning_effort (auto sends nothing; the rest map to OpenRouter's effort levels).
  • api_key - blank falls back to OPENROUTER_API_KEY (or legacy LLM_KEY); the node also reads a .env. Its tooltip admits the trade: a key typed here is saved with the workflow and can turn up in saved images and error reports. Use the env var unless you enjoy redacting.
  • zdr - off by default; enables OpenRouter's zero-data-retention routing.
  • regenerate - on by default, so every queue is a fresh paid call. Turn it off and ComfyUI's execution cache reuses the output while this node and its upstream inputs are unchanged. That's the in-memory cache, not a persistent one.

Nine media sockets grow in as you need them: image, image_2, image_3, and the same for video and audio.

Install

cd ComfyUI/custom_nodes
git clone https://github.com/theshubzworld/ComfyUI-OpenRouter-API.git
python -m pip install -r ComfyUI-OpenRouter-API/requirements.txt

Restart and search for Shubz or OpenRouter API; it lands under Shubz/OpenRouter. Dependencies are three packages - aiohttp, numpy, Pillow - so nothing fights your torch install and there are no model downloads. It does need ffmpeg and ffprobe on PATH once you connect a video or audio socket, plus ComfyUI 0.32.0+ and Python 3.10+. Then hand it a key:

export OPENROUTER_API_KEY="sk-or-v1-..."

Where people get burned

The preset default is the first. The second is key handling - a key in the widget travels with the PNG metadata and the workflow JSON you post when you ask for help.

A dropdown stuck on Loading OpenRouter models… or no compatible text-output model means the catalog hasn't been fetched yet, and that fetch runs inside the same timeout_seconds deadline as the call. A slow network plus a big video to compress can genuinely time out on run one - bump the timeout rather than deciding the node is broken.

On free models, expect rate limits. The :free variants are the tempting pick (listed first for exactly that reason), but the small ones get throttled hard, and the usual experience is a retry or two clearing it. Don't read three-per-modality as a provider promise either: OpenRouter publishes no universal attachment cap, so the nine sockets are a UI and resource bound, and a provider can still reject your nine-item request.

Last thing: the paid POST is never retried automatically, so a run that dies mid-request leaves you an ambiguous charge, not a double charge.

CategoryShubz/OpenRouter

Inputs (21)

NameTypeDefaultDescription
modelCOMBO— choose a compatible OpenRouter model —467 options: — choose a compatible OpenRouter model —, apodex/apodex-1.1-mini:free, cohere/north-mini-code:free, dots-studio/dots-3-note-preview:free, google/gemma-4-26b-a4b-it:free, google/gemma-4-31b-it:free, +461
reasoning_effortCOMBOauto8 options: auto, none, minimal, low, medium, high, +2
timeout_secondsINT1201–3600—
temperatureFLOAT1.000–2—
max_tokensINT81921–1000000—
response_formatCOMBOtext2 options: text, json_object
zdrBOOLEANfalse—
regenerateBOOLEANtrue—
system_presetCOMBORef2VA6 options: Ref2VA, I2VA, RefAud2VA, T2VA, FL2VA, None
system_promptSTRING—
user_promptSTRINGDescribe the supplied context.—
imageoptIMAGE—
image_2optIMAGE—
image_3optIMAGE—
videooptVIDEO—
video_2optVIDEO—
video_3optVIDEO—
audiooptAUDIO—
audio_2optAUDIO—
audio_3optAUDIO—
api_keyoptSTRINGOpenRouter API key. Leave blank to use OPENROUTER_API_KEY or LLM_KEY. Keys entered here are saved with the workflow.

Outputs (3)

NameTypeDescription
textSTRING—
infoSTRING—
creditsSTRING—