ModelScope API
Caption and interrogate images without spending a byte of VRAM
- image
- output
A vision model in the graph, on someone else's GPU
ModelScope is Alibaba's model hub - the HuggingFace of the Chinese ecosystem - and it runs an OpenAI-compatible inference endpoint at api-inference.modelscope.cn/v1. This node is a thin client for it: you give it a prompt and optionally an image, it gives you back a string. Nothing runs on your card. No 7GB VLM download, no VRAM budgeting, no second model fighting your checkpoint for space.
That last part is the reason it exists. A local captioner means holding a VLM resident next to your diffusion model, which is why half the local-LLM nodes do an unload/reload dance between them. This one is the other road: zero local footprint, free daily quota per account, and an endpoint that may be less prudish than a Western API. It's the API path from the LLM-in-ComfyUI story, aimed at captioning and interrogation - describe an image, seed a prompt from a picture, or answer "what is even in this render" without paying for it in VRAM.
What it sends, exactly
Because the pack's README is one sentence long, I read the source. The node builds an OpenAI chat request with a single user message containing a text block and an image_url block, where the image is a base64 JPEG data URL - your IMAGE tensor, first frame only, re-encoded at quality 85. Then it calls chat.completions.create with your model, max_tokens and temperature, and returns choices[0].message.content.
Two details you'd otherwise find by surprise:
- It always sends an image. Leave the
imageinput unwired and the code fabricates a blank 64×64 white image so the request stays well-formed. So yes, it doubles as a plain text LLM - you're just always paying for a vision-shaped call. api_tokensis a list. The field splits on commas, semicolons and newlines, and the node tries each token in order until one succeeds. That's quota rotation: the free tier is metered per token per day, so you stack a few accounts and burn them in sequence. If every token fails you get a Chinese-language error naming the last failure - roughly "all ModelScope API tokens failed, last error: …". It's not broken, it's out of tokens.
The inputs that matter
api_tokens(required, multiline) - grab an access token from your ModelScope account page (the API-Inference/SDK token section). Paste as many as you like, one per line.prompt- write it yourself. There's no built-in "describe this image" default, and an empty prompt sends an empty string, which the model answers however it feels like. For captioning, spell out the instruction and format: "Describe this image in one paragraph. Start with the subject and composition, then lighting, then style."image- optional, and only frame 0 of a batch. Need every frame of a 40-image batch described? That's a loop, not this node.model(defaultQwen/Qwen3-VL-8B-Instruct) - a free-text string, not a dropdown, so the roster is whatever the endpoint serves. Swap in a bigger Qwen-VL for better descriptions and slower calls.max_tokens(100–4000),temperature(0.1–2),seed. Be honest with yourself aboutseed: the code only seeds NumPy and sends nothing to the API, so it does not make generations reproducible. At the default temperature 0.7 you'll get a different caption on every re-run.
Output is a single output string. Wire it into the text input of a CLIPTextEncode (right-click → Convert text to input) for prompt-from-image, or into a save-text node if you're building .txt caption files.
Install
ComfyUI Manager → search ComfyUI-Practical-Tools, or:
cd ComfyUI/custom_nodes
git clone https://github.com/wenchengxiang/ComfyUI-Practical-Tools
# restart ComfyUI
This node needs openai>=1.0.0, which is in the pack's requirements.txt. If it's missing the node still appears but throws at queue time with a Chinese "openai 库未安装, 请执行 pip install openai" message - that's your cue, not a crash. The other requirements (onnxruntime, nvidia-vfx) belong to other nodes in the pack, and since the loader imports each node file separately, a [WCX Nodes Error] for one file leaves the rest working.
Where people get burned
Your image leaves the machine - that's the mechanism, not a bug. A base64 JPEG of your render goes to an Alibaba-hosted endpoint, logged and moderated under their policy. Don't point this at client work or anything private you'd mind being on someone else's disk.
Read the source before you run any API node. This category has already been weaponized once - ComfyUI_LLMVISION shipped credential-stealing malware and ended in a federal prosecution - and a node that holds your key and phones home by design is the one place where a malicious call-out looks normal. The good news: this file is short and boring, and the only host it contacts is api-inference.modelscope.cn. It's an obscure pack with no community footprint, so checking that is on you rather than on a reputation.
Caption quality has a known ceiling. Every VLM in this class mixes up attribution when two subjects share a frame - who's wearing what, who's doing what - and a caption can read as "impressive" while not actually reproducing the source image if you feed it back to a generator. For a few hundred training images, auto-caption then audit. For twenty, do it by hand.
Re-encoding before the model sees it. Images are JPEG quality 85 by the time they leave, so this is the wrong tool for reading fine print or spotting subtle artifacts in a 4K frame.
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| api_tokens | STRING | — | |
| imageopt | IMAGE | — | |
| promptopt | STRING | — | |
| modelopt | STRING | Qwen/Qwen3-VL-8B-Instruct | — |
| max_tokensopt | INT | 1000100–4000 | — |
| temperatureopt | FLOAT | 0.70.1–2 | — |
| seedopt | INT | -1-1–2147483647 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| output | STRING | — |