Extensions/ComfyUI_ExternalAPI
ComfyUI Extension

ComfyUI_ExternalAPI

ComfyUI node package for BYOK (bring-your-own-key) access to external generative APIs: Gemini, Veo, and more via LiteLLM.

By Tenafli·Created 3 months ago·Updated about 5 hours ago· 1
andy-ratsirarson/ComfyUI_ExternalAPI
Nodes2
On cloudLocal install
Categoryexternal_api/image, external_api/video
Stars1
Updatedabout 5 hours ago
Readme

ComfyUI_ExternalAPI

BYOK (bring-your-own-key) ComfyUI nodes for generating video and images through external providers — Gemini (Veo, Omni Flash, Imagen, Nano Banana), OpenAI Sora, Azure, and RunwayML — via LiteLLM. No local GPU-hungry model required; you supply your own provider API key and pay that provider directly.

Status: text-to-video works for all providers; image-to-video (a reference image guiding the animation) works for Gemini (Veo and Omni Flash) only. A separate image-generation node exists for Gemini only (Imagen for text-to-image, Nano Banana for both text-to-image and image-to-image) — see Roadmap for the other providers.

Features

  • One node, every provider. Pick a provider from a dropdown; the node's fields update to show only that provider's real parameters.
  • BYOK. Your API key never touches this repo or any third party beyond the provider itself. Prefer setting it as a server-side environment variable; an optional in-node override exists for convenience (see Security note).
  • Provider-accurate parameters. Each provider exposes its own supported aspect ratios, resolution tiers, and duration range — no guessing which combination is valid.

Supported providers & models

| Provider | provider value | Models | Aspect ratios | Resolutions | Duration | |---|---|---|---|---|---| | Google Gemini (Veo + Omni Flash) | gemini | Fetched live from your account via google-genai (filtered to models Google's own listing says still support video generation); falls back to a built-in list of Veo models (veo-3.1-fast-generate-preview, veo-3.1-generate-preview, veo-3.1-generate-001, veo-3.1-lite-generate-preview, veo-2.0-generate-001) plus Omni Flash models (gemini-omni-1.1-flash, gemini-omni-flash-preview) if that fails | Veo: 16:9, 9:16 · Omni: 16:9, 9:16 | Veo: 720, 1080, plus 4K on Veo 3.1 (non-Lite) — Veo 2.0 is 720-only · Omni: 360, 720, 1080, 4K | Veo: 4–8s (1080/4K are locked to 8s by Google, not enforced by this dropdown) · Omni: 3–10s | | OpenAI (Sora) | openai | sora-2, sora-2-pro, sora-2-pro-high-res | 16:9, 9:16 | 1K, 2K | 4–12s* | | Azure OpenAI (Sora) | azure | sora-2, sora-2-pro, sora-2-pro-high-res | 16:9, 9:16 | 1K, 2K | 4–12s* | | RunwayML | runwayml | gen3a_turbo, gen4_turbo, gen4_aleph | 16:9, 9:16 | 1K, 2K | 5s or 10s |

* OpenAI/Azure's upper duration bound isn't published by the provider; 12s is a conservative default rather than a confirmed hard limit.

Gemini's model list is fetched live from google.genai.Client().models.list() at node-registration time when a Gemini API key is available on the ComfyUI server, so new Veo/Omni releases show up automatically without an update to this package. The single gemini provider's model dropdown lists both families together — picking a Veo model reveals Veo's settings, picking an Omni Flash model reveals Omni's.

Omni Flash is a different API shape than Veo, under the hood. It runs through litellm's request/response Interactions API (litellm.interactions.acreate) instead of the polling avideo_generation/avideo_status/avideo_content calls Veo and the other providers use, so there's no in-progress polling for it — the node just waits on the single call. Only one-shot text-to-video is wired up today; Omni Flash's headline features — image/video-conditioned generation and multi-turn conversational editing — aren't implemented (see Roadmap), so its models don't appear in the image-source provider dropdown. Very large outputs that Google delivers by URI instead of inline base64 aren't downloaded yet and will raise a clear error rather than silently failing.

Image-to-video (a reference image conditioning the result) is currently supported for Gemini (Veo and Omni Flash) only. The other providers are text-to-video only for now — see Roadmap. The models, aspect ratios, resolutions, and durations in the table above apply to both text- and image-to-video.

Supported image providers & models

A separate node, API Image Generate (BYOK), handles still images. Gemini only for now:

| Model family | Example model ids | Text-to-image | Image-to-image | Sizes | |---|---|---|---|---| | Imagen | imagen-4.0-generate-001, imagen-4.0-fast-generate-001, imagen-3.0-generate-002 | ✅ | ❌ (no edit endpoint) | 1024x1024, 1792x1024, 1024x1792 | | "Nano Banana" (Gemini's multimodal image models) | gemini-2.5-flash-image (Nano Banana), gemini-3-pro-image (Nano Banana Pro), gemini-3.1-flash-image (Nano Banana 2) | ✅ | ✅, including multiple reference images | 1024x1024, 1792x1024, 1024x1792 |

Imagen calls litellm's aimage_generation (Google's :predict endpoint) and can only generate from text — it has no image-editing capability in litellm at all, so Imagen models don't appear in the node's image-source dropdown. Nano Banana models use the same call for text-to-image, and litellm's aimage_edit (Google's :generateContent endpoint) for image-to-image — which natively accepts a list of images, so batching multiple images into reference_image (the same convention as Omni Flash video, above) sends all of them as references in one call.

Both model families always return images inline (base64), never a hosted URL — this is fetched from Google directly, no local model.

Requirements

  • ComfyUI with the V3 node API (comfy_api.latest).
  • Python dependencies (installed automatically, see requirements.txt):
    • litellm==1.100.0
    • google-genai==1.47.0
  • An API key from at least one of the providers above.

Installation

Via ComfyUI Manager: search for ComfyUI_ExternalAPI and install.

Manually:

cd ComfyUI/custom_nodes
git clone https://github.com/andy-ratsirarson/ComfyUI_ExternalAPI.git
cd ComfyUI_ExternalAPI
pip install -r requirements.txt

Restart ComfyUI afterward.

Setup: API keys

Set the environment variable(s) for whichever provider(s) you plan to use on the machine running the ComfyUI server, then restart ComfyUI.

| Provider | Environment variable(s) | |---|---| | Gemini | GEMINI_API_KEY (or GOOGLE_API_KEY) | | OpenAI | OPENAI_API_KEY | | Azure OpenAI | AZURE_API_KEY, AZURE_API_BASE (your endpoint), AZURE_API_VERSION (optional) | | RunwayML | RUNWAYML_API_SECRET (or RUNWAYML_API_KEY) |

Security note

The node also has an optional api_key input that overrides the environment variable for that run. Prefer the environment variable. Anything typed into api_key gets saved into the workflow JSON and into output file metadata — so it can leak if you share a workflow or an output file. Only use it for quick one-off testing.

Usage

  1. Add the node: right-click canvas → Add Node → external_api/video → API Video Generate (BYOK).
  2. Enter a text prompt.
  3. Under source, leave it on text for text-to-video (see Image-to-video below for the image option).
  4. Pick a provider — the node reveals that provider's model dropdown.
  5. Pick a model — the node reveals that model's settings: size (aspect ratio), resolution, seconds (duration), and any provider-specific extras (e.g. Gemini's person_generation/negative_prompt, RunwayML's seed).
  6. Run the graph. The node polls the provider until the video finishes, then outputs a VIDEO.

Optional inputs: api_key (see Security note), poll_interval (seconds between status checks, default 10), timeout (max seconds to wait, default 600).

Image-to-video (Gemini: Veo & Omni Flash)

To animate a reference image instead of generating from text alone:

  1. Set source to image. The node reveals a provider dropdown that currently lists only gemini.
  2. Pick gemini, then pick either a Veo or an Omni Flash model — the same combined list as text-to-video.
  3. The node reveals settings for that model: both show size, resolution, seconds, and reference_image; Veo additionally shows person_generation and an optional negative_prompt (Omni doesn't have either — it only takes text + image + video content).
  4. Connect reference_image — an IMAGE socket, same as any other ComfyUI node that takes an image. Wire in a Load Image node to upload/pick a file, or connect any other node's IMAGE output directly (including an image generated earlier in the same workflow — no need to save it to disk first). This input is technically optional at the graph level, but image-to-video will fail at runtime with a clear error if left unconnected.
  5. For multiple reference images (Omni Flash only): feed a batch into reference_image instead of a single image — e.g. chain several images through a Batch Images node upstream. Omni Flash sends each image in the batch, in order, as its own reference alongside the prompt. Veo does not support this: if it receives a batch of more than one image it raises a clear error rather than silently using just the first (litellm's Veo integration only has a single-image field — see Roadmap).
  6. Enter a prompt. It's still required with image-to-video — the prompt directs how the reference image(s) are animated.
  7. Run the graph. Veo polls until the video finishes; Omni Flash is a single request with no status-polling endpoint, so its progress bar just ticks up with elapsed time rather than tracking real progress — both output a VIDEO.

Note: for Veo image-to-video, person_generation only offers allow_adult (text-to-video also offers allow_all). Google restricts generation to adults when a reference photo is provided, so the broader option isn't available here.

Veo is single-image only; Omni Flash accepts multiple, per above. Neither supports tagging/naming individual images for reference within the prompt text (Omni Flash images are just referenced positionally/by natural language, e.g. "the first image" — there's no @name syntax on Google's side), video input (Omni Flash also accepts a short video, not just images), or Veo's interpolation/video-extension — see Roadmap.

Image generation (Gemini: Imagen & Nano Banana)

  1. Add the node: right-click canvas → Add Node → external_api/image → API Image Generate (BYOK).
  2. Enter a text prompt.
  3. Under source, leave it on text for text-to-image, or set it to image for image-to-image (see below).
  4. Pick provider → gemini, then a model — Imagen or Nano Banana for text; Nano Banana only for image (Imagen has no edit capability).
  5. The node reveals size (1024x1024 / 1792x1024 / 1024x1792) and num_images (1–4); the image source also reveals reference_image.
  6. For image-to-image, connect reference_image the same way as the video node's — a single image, or a batch (via a Batch Images node) for multiple references, all sent to the same Nano Banana call.
  7. Run the graph. This is a single request with no polling (image generation is fast) — it outputs an IMAGE, connect it to any node that takes one (e.g. Save Image, Preview Image, or straight into another generation node).

Optional input: api_key (see Security note).

Roadmap

  • Image generation for other providers — implemented for Gemini only (Imagen + Nano Banana; see above). OpenAI (DALL-E 3, gpt-image-1), Azure, and RunwayML (gen4_image/gen4_image_turbo, currently unused since they were miscategorized as video models — see below) are future work.
  • RunwayML gen4_image/gen4_image_turbo miscategorization — these are image models, not video, but were previously listed (and still callable) under the video node's RunwayML provider; they've been removed from there but not yet added to the image node.
  • Image-to-video — implemented for Gemini (both Veo and Omni Flash; set source to image). Support for the other providers (OpenAI/Azure Sora, RunwayML) is future work.
  • Veo multi-image references — Veo 3.1 itself supports up to 3 reference images plus last-frame interpolation and video extension, but litellm's Veo integration (avideo_generation) only exposes a single image field with no way to pass more — reaching Veo's real multi-image capability would mean bypassing litellm with a raw HTTP call, which isn't done here. Omni Flash, by contrast, already accepts a batch of any size (see above) since its input is just a list we build ourselves.
  • Gemini Omni Flash: video input and conversational editing — Omni Flash also accepts a short video as conditioning input (not just images) and supports iteratively refining a result across turns via previous_interaction_id. Neither is wired up yet.
  • Gemini Omni Flash: large (uri-delivered) outputs — Google delivers big videos by URI instead of inline base64; downloading those isn't implemented yet, so such a response raises a clear error instead of returning a video.

Development

pip install -e ".[dev]"
pytest

Tests fake just enough of comfy_api.latest (see tests/conftest.py) to exercise the nodes outside a real ComfyUI install — no GPU or ComfyUI checkout required.

License

MIT — see LICENSE.