ComfyUI Node
Image to Text
A ComfyUI node in Transformers/Multimodal/ImageToText with 3 inputs and 1 output.
Image to Text
- image
- generated_text
◄model_nameSalesforce/blip-image-captioning-base►
◄max_new_tokens50►
CategoryTransformers/Multimodal/ImageToText
Inputs (3)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | — | |
| model_name | STRING | Salesforce/blip-image-captioning-base | — |
| max_new_tokens | INT | 501–512 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| generated_text | STRING | — |