ComfyUI Node
π± Gemini Vision
Gemini Vision node outputs a rich description of the input image.
π± Gemini Vision
- image
- system_instruction
- response
βtext_promptDescribe this image in detail.βΊ
βapi_keyβΊ
βmodelgemini-2.5-flashβΊ
βmax_tokens5000βΊ
βtemperature0.7βΊ
CategoryArtha/LLM/GEMINI
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | β | |
| text_prompt | STRING | Describe this image in detail. | β |
| api_key | STRING | API key will be visible in plain text. Consider adding your api to the api.json located inside this custom node folder. | |
| model | COMBO | gemini-2.5-flash | 5 options: gemini-2.5-pro, gemini-2.5-flash, gemini-2.5-flash-lite, gemini-2.0-flash, gemini-2.0-flash-lite |
| max_tokens | INT | 50001β8192 | For Gemini models, a token is equivalent to about 4 characters. 100 tokens is equal to about 60-80 English words. |
| temperature | FLOAT | 0.70β2 | A temperature of 0 means only the most likely tokens are selected, and there's no randomness. Conversely, a high temperature injects a high degree of randomness into the tokens selected by the model, leading to more unexpected, surprising model responses. |
| system_instructionopt | ARTHAINSTRUCT | β |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| response | STRING | β |