DIGIT GPT Image
GPT Image 2 in ComfyUI, with a real inpainting mask
- image1
- image2
- image3
- image4
- image5
- image6
- image7
- image8
- image9
- image10
- image11
- image12
- image13
- image14
- image15
- image16
- mask
- image
- text
GPT Image 2 is OpenAI's image model, and it doesn't run on your GPU - it runs on OpenAI's servers, which is why this node goes through fal.ai instead. You pay per image with a FAL_KEY, and the whole thing behaves like a ComfyUI generator node: prompt in, IMAGE tensor out, ready to feed the rest of your graph. The part that's genuinely nice here is the edit mode, which gives you a real mask input for inpainting - white areas get edited, black stays put.
The mode auto-detects from what you connect. No images → text-to-image. Connect image1 through image16 → edit mode, and add a mask and it's inpainting. Same node, different job, no mode switch to forget.
How it works
The node talks to fal's hosted openai/gpt-image-2 endpoint. Your prompt (and any images) go up, the job runs on fal's infrastructure, and the result comes back as an IMAGE tensor. The mechanism is dead simple from the node's side; the complexity is all in the API.
The inputs that matter:
- prompt - required. GPT Image reads natural language well, so write a sentence, not a tag list.
- model -
gpt-image-2, the only one. No dropdown variety here. - image_size -
auto,square_hd,square,portrait_4_3,portrait_16_9,landscape_4_3,landscape_16_9, orcustom. Pickcustomto unlockcustom_width/custom_height(multiples of 16, max edge 3840). - quality -
auto/low/medium/high. Higher quality costs more per image; the tooltip says so flat out. - output_format -
png,jpeg, orwebp. - num_images - 1–4 images per API request. Total output = num_images × batch_count.
- batch_count - 1–128 parallel fal jobs, returned as one IMAGE batch. This is where the bills start getting interesting.
- seed - read the tooltip: "Re-run control only - the GPT Image API has no seed. 0 regenerates on every queue." Don't expect reproducibility you don't have.
- image1…image16 - the edit inputs. One image for a simple edit; several for multi-image work.
- mask - inpainting mask. White is edited, black is preserved. This is the underrated one; most API wrappers don't expose a real mask.
Outputs are image (the IMAGE batch) and text (usually empty, but the node returns it alongside). The built-in resilience retries up to three times - except content-policy errors (HTTP 422), which are not retried because retrying won't change the answer.
Install
cd ComfyUI/custom_nodes
git clone https://github.com/thedepartmentofexternalservices/comfyui-digit.git
cd comfyui-digit
pip install -r requirements.txt
export FAL_KEY=your_fal_key
(Or ComfyUI Manager → search comfyui-digit → install.) Restart ComfyUI, look under DIGIT.
Common issues
The two failure classes are the key and the filter. No FAL_KEY set → the node can't authenticate, so export it before launching ComfyUI. Content-policy 422s mean OpenAI refused the prompt - GPT Image has OpenAI's moderation welded in, and there's nothing the node can do about it. That's the "filter follows the model, not the node" reality of closed APIs.
Cost is the one people underestimate. GPT Image at high quality, large size, num_images 4, batch_count 8, is 32 API calls in one queue. It's per-image metered and it adds up in a way local generation never does. Use quality: auto for drafts and only spend high when a frame actually matters.
Inputs (27)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | — | |
| model | COMBO | gpt-image-2 | 1 options: gpt-image-2 |
| image_size | COMBO | auto | Preset size, or custom to use custom_width/custom_height. |
| quality | COMBO | high | Higher quality and larger sizes cost more per image. |
| output_format | COMBO | png | 3 options: png, jpeg, webp |
| num_images | INT | 11–4 | Images per API request. Total output = num_images x batch_count. |
| seed | INT | 00–2147483647 | Re-run control only — the GPT Image API has no seed. 0 regenerates on every queue. |
| image1opt | IMAGE | — | |
| image2opt | IMAGE | — | |
| image3opt | IMAGE | — | |
| image4opt | IMAGE | — | |
| image5opt | IMAGE | — | |
| image6opt | IMAGE | — | |
| image7opt | IMAGE | — | |
| image8opt | IMAGE | — | |
| image9opt | IMAGE | — | |
| image10opt | IMAGE | — | |
| image11opt | IMAGE | — | |
| image12opt | IMAGE | — | |
| image13opt | IMAGE | — | |
| image14opt | IMAGE | — | |
| image15opt | IMAGE | — | |
| image16opt | IMAGE | — | |
| maskopt | MASK | Optional inpainting mask (edit mode only). White areas are edited. | |
| custom_widthopt | INT | 1024320–3840 | Used only when image_size is custom. Multiples of 16. |
| custom_heightopt | INT | 1024320–3840 | Used only when image_size is custom. Multiples of 16. |
| batch_countopt | INT | 11–128 | Number of parallel fal jobs. Each is a separate API call; results return as one IMAGE batch. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| text | STRING | — |