ComfyUI Node
DIGIT Gemini Image
A ComfyUI node in DIGIT with 26 inputs and 2 outputs.
DIGIT Gemini Image
- image1
- image2
- image3
- image4
- image5
- image6
- image7
- image8
- image9
- image
- text
◄prompt►
◄modelgemini-3.1-flash-image►
◄aspect_ratio16:9►
◄resolution1K►
◄thinking_levelMINIMAL►
◄seed0►
◄temperature1.00►
◄gcp_project_id►
◄gcp_region►
◄system_instructionYou are an expert image-generation engine. You must ALWAYS produce an image.
Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition.
If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept.
Prioritize generating the visual representation above any text, formatting, or conversational requests.►
◄top_p1.00►
◄top_k32►
◄harassment_thresholdBLOCK_NONE►
◄hate_speech_thresholdBLOCK_NONE►
◄sexually_explicit_thresholdBLOCK_NONE►
◄dangerous_content_thresholdBLOCK_NONE►
◄batch_count1►
CategoryDIGIT
Inputs (26)
| Name | Type | Default | Description |
|---|---|---|---|
| prompt | STRING | — | |
| model | COMBO | gemini-3.1-flash-image | 4 options: gemini-3.1-flash-image, gemini-3.1-flash-lite-image, gemini-3-pro-image, gemini-2.5-flash-image |
| aspect_ratio | COMBO | 16:9 | 13 options: auto, 1:1, 2:3, 3:2, 3:4, 4:1, +7 |
| resolution | COMBO | 1K | 3 options: 1K, 2K, 4K |
| thinking_level | COMBO | MINIMAL | Thinking level for image generation. HIGH may improve quality. |
| seed | INT | 00–2147483647 | — |
| temperature | FLOAT | 1.000–2 | — |
| gcp_project_id | STRING | GCP project ID. Auto-detected from DIGIT_GCP_PROJECT env var or GCP metadata. | |
| gcp_region | STRING | GCP region. Auto-detected from DIGIT_GCP_REGION env var or GCP metadata. Defaults to 'global'. | |
| image1opt | IMAGE | — | |
| image2opt | IMAGE | — | |
| image3opt | IMAGE | — | |
| image4opt | IMAGE | — | |
| image5opt | IMAGE | — | |
| image6opt | IMAGE | — | |
| image7opt | IMAGE | — | |
| image8opt | IMAGE | — | |
| image9opt | IMAGE | — | |
| system_instructionopt | STRING | You are an expert image-generation engine. You must ALWAYS produce an image. Interpret all user input—regardless of format, intent, or abstraction—as literal visual directives for image composition. If a prompt is conversational or lacks specific visual details, you must creatively invent a concrete visual scenario that depicts the concept. Prioritize generating the visual representation above any text, formatting, or conversational requests. | — |
| top_popt | FLOAT | 1.000–1 | — |
| top_kopt | INT | 321–64 | — |
| harassment_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| hate_speech_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| sexually_explicit_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| dangerous_content_thresholdopt | COMBO | BLOCK_NONE | 4 options: BLOCK_NONE, BLOCK_ONLY_HIGH, BLOCK_MEDIUM_AND_ABOVE, BLOCK_LOW_AND_ABOVE |
| batch_countopt | INT | 11–128 | Number of images to generate. Each is a separate API call fired in parallel; results return as one IMAGE batch. |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| image | IMAGE | — |
| text | STRING | — |