Runware.ai ComfyUI Inference API Integration
Runware Inference API Integration for ComfyUI (No GPU Required).
Nodes (380)
Write a whole song from a prompt — ACE-Step makes lyrics, key, and BPM first-class inputs
A music model that lets you set the BPM, key, and time signature
The music node for people who actually want to steer the sampler
The full-quality ACE-Step, for the take that matters
Full songs from a text prompt, fast
Alibaba's video editor, prompt your existing clips
Alibaba's sleeper video model with the size presets
Ask a vision model about your image, get a string back
The image-to-text node that reads your images
Clone a voice from a 3-second clip — Qwen3-TTS as a ComfyUI node
Give Your Video a Voice It Didn't Have
Design a voice with a sentence — no reference clip required
The open editing model, minus the GPU
The unified 7B, API-only for a reason
Qwen-Image 2.0 Pro — the API-only refresh of the open image flagship
Alibaba's text-rendering flagship in the cloud
Describe the change, skip the mask
The better-behaved editor
Compositing layers out of thin air (PNG-keepers, this is for you)
The Wan That Never Got Open Weights, Rendered
Wan2.5-Preview Image — a Wan nobody can download, used through Runware
Wan 2.6 without a 27B model on your desk
Image-to-video with native audio, served from the cloud
Alibaba's image model with a surprising resolution ceiling
Wan2.7 — the API-only Wan you couldn't run locally anyway
Wan 2.7 Image — the newest closed Wan, with a thinking-mode switch
The Wan You Can Actually Run
Z-Image in the cloud — the 'SDXL 2.0' you can also afford to skip local for
The 6B model that killed the FLUX 2 hype
Claude Fable 5 — extended thinking with built-in prompt caching
Claude Haiku 4.5 in a node — the fast, cheap LLM for graph plumbing
Claude Opus 4.7 — the reasoning tier for when the graph has to think
Claude Opus 4.8 — the newest Opus, and the priciest node in the pack
Claude Sonnet 4.6 — the middle child that usually wins
Run an illustration checkpoint without owning a GPU
FLUX.1 [dev] and every FLUX.1 fine-tune, without the 24GB download
Kontext dev's edits, without the 12B-model hardware bill
Any FLUX checkpoint, zero downloads
Danbooru tags, no VRAM
NoobAI XL without the vpred setup headache
Every Pony checkpoint, one node
Any CivitAI SD 1.5 checkpoint, on Runware's GPUs
Run any SDXL checkpoint without downloading a byte
The fast SDXL rung, on any checkpoint
SDXL Turbo in the cloud, with the whole SDXL ecosystem attached
Z-Image — but any fine-tune you want, by AIR
Z Image Turbo, but bring your own checkpoint
The text-in-image specialist that still lost
The 8-step shortcut to text-in-image
BFL's closed flagship, on your canvas
Nano Banana, as a ComfyUI node
FLUX.1 [dev] without the 24GB VRAM tax
Growing your canvas, credibly
The Swiss-army editor
Inpainting without the VAE churn
The Edit-By-Instruction Model, In the Cloud
The API-tier instruction editor
Same Editing Magic, Fewer Knobs, Hosted
The fast, Apache-licensed FLUX
FLUX.2 [dev] without the 18–24GB VRAM bill
The Cheap, Fast Tier of the New Flux
The Apache-2.0 FLUX that runs on a lunch budget
The size-distilled Flux 2, with actual dials
The Editing Daily Driver, Without the 21GB
The distilled Flux 2 that fixed the VRAM wall
FLUX.2 Klein 9B — the consumer-grade Flux, served cloud-side
The 32B flagship that doesn't need your VRAM
The Flagship, Where Everything Is Decided For You
Remove anything, and let the background heal
Extend the Canvas, Keep the Subject
Person, garment, and a 4-step render
The benchmark-bred edge specialist
Hair-level background removal, no VRAM spent
The cheap bulk cutout machine
When the fine detail lives at high resolution
The 'trained on everything' cutout model
The hair-saving cutout, minus the model download
Background removal that survives hair
The background remover people actually reach for
The camouflage hunter you'll rarely need
The commercially-safe generator with the moderation trap
Bria FIBO without the 16GB download — cloud text-to-image that's built for editing
Iterative image editing the way FIBO was meant to be used
Seven image editors in one node, and one of them is a time machine
Bria FIBO Lite — the version of FIBO that took away your dials
The one-knob cloud upscaler — 2x or 4x, zero VRAM
Product-photo background swaps without cutting out anything by hand
The zero-settings background remover — Bria RMBG v2.0, one input, one output
RMBG v2.0 Open — the same cutout, under Runware's own catalog
Video background removal, and the one thing that will trip you first
Object erasure in video — make the thing disappear, keep the audio
2x or 4x video upscaling on the cloud — with a frames-not-video twist
The speed levers — TeaCache, DeepCache, and friends for cloud inference
Audio Settings — the builder with three knobs and a lot of defaults
Cloud ControlNet — spatial structure, minus the local model files
Embeddings — textual inversion for Runware models, if you've got one
Import Model — how your own checkpoint ends up on Runware's GPUs
The IP-Adapter you don't have to install
The builder node that makes 'lora' sockets usable
Messages — build a conversation, one node at a time
Outpaint — extend the canvas, let the model invent what's beyond the edge
PhotoMaker — keep the face, change everything else
PuLID in the cloud — face identity without the InsightFace install
Reference Images — the @-tag system that powers try-on and lip-sync
Feed a Clip Into a Video Model
Clone a Voice From a Clip — and Please, Get the Transcript Right
Handing the Last Steps to a Refiner
The Builder That Feeds Every TTS Model
YOLO Finds the Faces, Diffusion Fixes Them
One Photo, One Audio Clip, One Animated Human
Same Party Trick, Now With a Mask and a Fast Lane
Text-to-Video That Respects Your Camera
The Iteration Mode of the Same Video Model
Draft Cheap, Promote the Winner, Ship With Sound
The full-quality ByteDance video model, in the graph
The speed tier of ByteDance's video model
ByteDance's video model as a ComfyUI node
One Prompt for Voices, SFX, Music, and Dialogue
ByteDance's Flagship Image Model, Without the Weights Problem
ByteDance's flagship, minus the API plumbing
ByteDance's image model, fast tier, sequence mode
ByteDance's Seedream 5.0 Pro — great text rendering, huge canvas, zero local VRAM
Crank a video to 4K in the cloud, BytePlus edition
BytePlus video enhancement with a style dial for the source
The Video Upscaler That Returns an Image, and Other Schema Quirks
CCSR — the detail-restoring upscaler, without the 12GB+ VRAM tax
Runware's Own Upscaler, Prompt Included
A pre-flight NSFW check that costs pennies
The edge-map factory for cloud ControlNet
The Preprocessor That Sits Before Your ControlNet
The Character-Consistency Shortcut
Clean Lines for Cleaner Characters
Straight-line ControlNet conditions, without the annotator stack
ControlNet for People Who Care About Light
Make Your Character Do the Reference's Move
Turn a doodle into a ControlNet condition
Control layout with a segmentation map, not a sketch
Steal an Image's Ingredients, Not Its Composition
The Forgiving ControlNet Preprocessor
The tile preprocessor that powers cloud upscale workflows
Feed it a still and a soundtrack, get a clip
The same audio-to-video, on a budget clock
The escape hatch for models the pack hasn't met yet
An LLM inside your ComfyUI graph
An LLM Living Inside Your ComfyUI Graph
The bigger, fiddlier TTS model for when you want control
A LoRA-style tune-up that rents the GPU
Reference-driven photos, zero prompt required
Reference in, out-of-focus world out
For the shots that snap
The Film-Grain Node That Doesn't Even Have a Prompt Box
The all-rounder of the reference-photo family
Golden-hour vibes, from a node with no prompt
Give your ComfyUI workflow a voice — cloud TTS that can clone one
The dev-tuned FLUX that ditched the plastic look
Pull one field out of a JSON firehose
The cheapest LLM in the graph
Google-grade speech, straight into the graph
The thinking model for the hard parts of your workflow
The one that searches the web and answers in JSON
An LLM in your graph, no API account of your own
Google's Video Model That Takes In Everything
A 31B reasoning model inside your ComfyUI graph — with no GPU of your own
Google's image gen, minus the local GPU
The speed-quality hybrid that finally takes video refs
Nano Banana 2 Lite is Google's image model inside your ComfyUI graph
Google's Gemini image model, no Google account needed
Veo 3.1 — the frontier video model — from a ComfyUI node
Veo 3.1 Fast drops Google's best video into a ComfyUI node — audio included
Google's Best Video Model, on a Budget
HeyGen Avatar IV turns a still photo into a talking presenter
A presenter, on demand, straight from a dropdown
A talking-head avatar in ComfyUI — 1,288 avatars on tap
HiDream-I1 Dev — the MIT-licensed prompt-follower, minus the 17B weight problem
The prompt-adherence specialist, at iteration speed
HiDream-I1 at full precision, no 40GB download
Rodin Gen-2 puts a 3D modeler in your ComfyUI graph
The text-on-image specialist, parked inside your graph
Ideogram 2.0 Remix edits your image with a strength dial
Ideogram 2a is the lean cloud option for text-on-image
The lightweight remix node that keeps your text
Character consistency plus 63 style presets in one node
Ideogram 3.0 Edit
Ideogram 3.0 Reframe
Ideogram 3.0 Remix
Product shots without the photo studio
The text-in-image champ, with its own prompt schema
The text-in-image king, without the 24GB download
The hosted Ideogram 4 remix — no weights, no filter fights
Pull text out of an image like it's a layer
Mask it, delete it, keep the rest pixel-identical
The cheap-and-cheerful text-to-image node
ImagineArt 1.5 Pro
ImagineArt 2.0
Voice that doesn't sound like it's reading a spec sheet
The fast, cheap voice for first passes
Realtime voice, minus the sliders
The Runware node that actually feels like local ComfyUI
The cloud video model that hides its GPU from you
The value tier for testing cloud video cheaply
Image-to-video that finally feels like a production tool
Cheaper I2V when the clip isn't the hero
The generation people actually talked about
The 1.6 look on a value-tier budget
The ceiling the open-source world was chasing
The flagship that fixed the details
Flagship image-to-video with the frame wired in
The flagship's cheaper image-to-video half
Speed as a feature, motion as the product
Kling 2.5 Turbo on your ComfyUI canvas, without the 80GB checkpoint
A talking head from one photo and an audio file
Talking-head video at a price you can loop
Kuaishou's image model, on your ComfyUI canvas
The reference-driven image model with 'from input' sizes
The one that wants your exact pixel dimensions
New audio for an existing video, done in the cloud
The everything-input video node with sound
Reference-driven video, trimmed for cost
15-second clips at real resolution, with sound
Reference-video editing at cinema resolution, no GPU required
The middle tier that most reference-video edits deserve
The cheap way to iterate on reference-video edits
The text-to-video workhorse, with a negative prompt and real controls
The budget text-to-video node for prompt roulette
Kling 3.0 Turbo's multi-shot prompt is the hidden superpower most people miss
The 'thinking' tier for edits that need more than a prompt pass
The cheap lane for learning what O1 edits can do
The hype model, served hot via API with no 24GB VAE hunt
The cheaper tier for testing prompts against the hype model
The hosted Krea 2 that follows prompts better than the open weights
The undistilled base, served over an API with real sampler dials
The 12B base everyone's building on
The Aesthetic Flux, With Every Socket Attached
The synchronized-audio video model, now a single ComfyUI node
The release that made LTX competitive, now without the local VRAM tax
The distilled tier that trades steps for speed, without losing the LoRA socket
The presets-only tier that makes one question go away
The audio-synced video model that Lightricks bet the family on
Fix one bad segment of a video, not the whole clip — LTX-2 Retake, the cloud way
A cheap way to stop writing bad prompts yourself
The tiny node that feeds every other Runware node
The reframe node for people who think in canvases, not clips
The image model that will Google for you before it draws
The bigger-budget sibling for when the first pass isn't enough
A cloud mask maker for the face-fix pipeline
Want a clean mask of just the eyes? This cloud MediaPipe node hands you one
The License-Clean Face Mask You've Been Missing
Full-face masks in one click, from the cloud
An eyes-only face mask, via the API — the boring node that earns its keep
The tighter face mask, for edits that shouldn't wander
A lips-only mask, without drawing a single pixel
One mask for the eyes and nose — the 'what are you looking at' region
Nose + lips in one mask, for the middle of the face
A nose-only mask — yes, this is a real use case
Ask a model how old the person in your video is
Auto-caption a video in one node — no local vision model
From a photo to a textured GLB, without touching Blender
A vision-language model that talks about your image, not just describes it
The 'turn this one object into a 3D model' node
Image to a real, textured 3D mesh — without the 3D stack
MiniMax 01 — the cloud video model with a prompt optimizer on tap
When your prompt is a shot list
The Video Node With No Required Inputs At All
The earlier MiniMax video tier that still earns its spot
Cloud video gen that isn't a local-VRAM nightmare
The image-to-video express lane
Yes, this pack does text LLMs too
A long-context LLM in the graph
MiniMax M3 without a GPU
Full songs from a prompt in ComfyUI, no GPU, no model files
Restyle an existing song into a cover version, in-graph
Character voices and dialogue without a TTS model download
Type 'door creak' and get actual door-creak audio
Sound effects from a prompt, because you shouldn't have to record a door
Kimi K2.6 — a 1M-token-reasoning model as a ComfyUI node
The clean way to delete things from a photo
A one-input node that guesses age from a face
A caption for any image, from the text encoder that ran SD 1.5
A full-featured GPT-5.4 chat node, straight into ComfyUI
The everyday GPT-5.4 — fast, cheap, same controls
When the job is tiny and you're calling it a lot
The heavy thinker with the training wheels removed
The newest model in the pack, JSON-native and thinking by default
The compact model that still speaks JSON
The Runware text node
The text-rendering king, without the OpenAI subscription
Same socket, sharper output
The cheap seat on the text-rendering train
GPT Image 2 in your ComfyUI graph — with moderation dialed to 'low'
OpenAI's video model without the API plumbing
The 'turn the dial up' Sora node
Turn raster art into a real SVG
Make any video talk, and mean it
Edit an existing video, not just make one
The workhorse text-to-video, effects included
The camera gets a director
Where the TikTok effect templates live
The quality pick of the family
The version where PixVerse got audio and 'thinking'
The thinking, sound-on refresh
Iterate like you mean it
Hosted video gen in your graph
Pruna's take on text-to-image, minus the GPU
The image editor that keeps the prompt simple
A fitting room in a node
Outsource the upscale, keep the quality
Text-to-video with the sound already baked in
Bring a reference to life
Make a still image talk — P-Video-Avatar turns one face into a lip-synced clip
Swap the subject, keep the performance
How old is this person, for real? Runware's Qwen2.5-VL age check
The old reliable upscaler, now with no VRAM at all
The designer's model, now a plain ComfyUI node
The refresh, and when it's actually worth it
The better rendering, at a per-image price
The boring Recraft that does real work
The pristine product shot, on demand
The flagship design model, one node away
Prompt in, honest-to-goodness SVG out
The prompt-to-SVG node that just works
Turn the raster you already have into an SVG
Cutouts without the model downloads
A closed 4K photorealism model you can reach from a ComfyUI node
The realism brand, moved to Flux and the cloud
Near-instant realism, no GPU
Juggernaut Pro Flux, no 24GB VRAM required
The photorealism default, now on Z-Image
Train a FLUX.1 style LoRA on someone else's GPU
Train a FLUX.2 Klein 4B style LoRA without touching your GPU
FLUX.2 Klein 9B style LoRA training, with the VRAM math handled for you
Train a Qwen-Image style LoRA in the cloud — your GPU gets the day off
Train a Z-Image style LoRA from ComfyUI — no GPU, no Kohya
Video-to-video from the company that made SD 1.5
Runway Aleph 2.0 video edits, from inside your graph
The image-to-video node that only needs frames
Runway Gen-4 images, without the Runway website
Reference-driven images on the fast lane
Frames in, motion out, in ComfyUI
SkyReels V4 — the video model that can add its own sound
The design model the internet is sleeping on
Riverflow 2.0 Pro — the cloud image node with built-in upscaling reference
Thinking, transparency, and a scoring rubric
The 4K tier with an extra level of thinking
A real Stable Diffusion 3 node that runs on someone else's GPU
The boring checkpoint that saves you a headache
The classic latent upscaler, now with nothing to download
The Cloud Upscaler That Skips the Whole Local Dance
Make a video's mouth match any audio, without running a lip-sync model locally
A lipsync node that actually takes your audio
Make the mouth match the audio, without a local model
Re-animate a face to a new performance — the deepfake-adjacent node
Generating .glb files from your ComfyUI graph
Prompt-to-GLB without owning a 24GB GPU
Topaz's generative video upscaler, as a cloud ComfyUI node
A GLB out of a text prompt (or an image)
The local-realism favorite, hosted and fast
The boring node that makes the whole cloud pack work
Turn a photo and an audio clip into a talking presenter, in the cloud
Vidu 1.5 — the video node with a BGM switch and movement dial
The pragmatic text-to-video pick on Runware
The style-preset video model that hides its prompt box behind the fun part
The image-to-video node with no prompt box — and that's the point
The odd one out — a Vidu image generator that needs a reference photo
The generation-two Vidu that finally lets you dial in fractional seconds
The fast sibling of the Vidu Q2 line, prompt required
The Chinese studio's video model, no account required
The newest Vidu, with native audio and a demand for explicit dimensions
A one-socket 'how old is this person' node
A full LLM node that needs a messages builder to talk to
XAI's generator as a plain ComfyUI node
The least-censored image model you can wire straight into ComfyUI
XAI's Video Model, Now a ComfyUI Node
Turn a frame into motion with xAI's Grok Imagine Video
The pack's shared speech engine, exposed as a bare audio node
The tiny detector that feeds your face workflows
Cloud hand detection that hands you a mask, not a bounding box
Person silhouettes as masks, no local segmentation stack
A face mask in seconds, for fix-the-face workflows without the local model
The small-model person segmenter for when nano isn't cutting it
Z.ai's newest model, tool-calling enabled, with a bigger token budget
ComfyUI-Runware
Every Runware model as a ComfyUI node: image, video, audio, 3D, and text, all running in the cloud. No local GPU and no per-model setup. The whole catalog shows up in your node menu, and each node's widgets are the model's real parameters.
Full guide: https://runware.ai/docs/platform/comfyui
Install
ComfyUI Manager (recommended): open the Custom Nodes Manager, search Runware, install, and restart.
Manual:
cd ComfyUI/custom_nodes
git clone https://github.com/Runware/ComfyUI-Runware
pip install -r ComfyUI-Runware/requirements.txt
Restart ComfyUI.
API key
Create a key in the dashboard, then provide it one of these ways:
- ComfyUI Settings → Runware API key: paste it in the UI, no terminal needed.
RUNWARE_API_KEYenvironment variable: overrides the Settings field, useful for servers.- Runware CLI: run
runware auth loginonce and the nodes reuse the stored key.
Quick start
- Double-click the canvas and search a model by name (e.g.
FLUX.2 [dev]), or browseRunware/Image. - Type your
positivePromptand set the dimensions. - Wire the node's IMAGE output into Preview Image or Save Image.
- Queue. The request runs on Runware and comes back as a native
IMAGE.
What's in the pack
- One node per model, grouped
Runware/<Modality>/<creator>. Widgets are the model's real parameters, with correct ranges, defaults, and dropdowns. - Native outputs: image, upscale, and background-removal return
IMAGE; audio returnsAUDIO; video returnsVIDEO; 3D and other files save to your output folder and return a path; text returns a string. - Native inputs: reference and seed images are
IMAGE, inpainting masks areMASK; audio, video, and document inputs take a URL, path, or UUID. - Builder nodes (
Runware/Params) keep model nodes clean. Stackable features (LoRA, ControlNet, IP-Adapter, Embeddings, and more) each wire into a model's typed socket, and you chain the stackable ones to combine them. - Custom checkpoints:
Runware/Custom modelshas a node per architecture (SDXL, SD 1.5, FLUX, Pony, and more) for community fine-tunes. - New models, day one:
Runware (custom)takes any model AIR andtaskTypeas JSON and returns raw JSON. Pair it withRunware Getto pull a field out of the response (e.g.0.imageURL). - Run info on the title bar: each run shows its cost and, when a content check ran, the NSFW result (e.g.
$0.00078 · NSFW: no).
The full guide covers every builder, custom models, parameter behavior, and troubleshooting.