ComfyUI Node
Qwen2_5_VL_ImageToTextpose
A ComfyUI node in Qwen2.5-VL with 7 inputs and 2 outputs.
Qwen2_5_VL_ImageToTextpose
- image1
- image2
- raw_description
- cleaned_description
◄max_new_tokens2048►
◄max_description_length1002►
◄model_pathhelenai/Qwen2.5-VL-7B-Instruct-ov-int4►
◄deviceCPU►
◄styleNone►
CategoryQwen2.5-VL
Inputs (7)
| Name | Type | Default | Description |
|---|---|---|---|
| image1 | IMAGE | — | |
| max_new_tokens | INT | 20481–2048 | — |
| max_description_length | INT | 100250–2048 | — |
| model_path | STRING | helenai/Qwen2.5-VL-7B-Instruct-ov-int4 | — |
| device | COMBO | CPU | 2 options: CPU, GPU |
| style | COMBO | None | 7 options: Standing, sitting, walking, t_pose, running, squatting, +1 |
| image2opt | IMAGE | — |
Outputs (2)
| Name | Type | Description |
|---|---|---|
| raw_description | STRING | — |
| cleaned_description | STRING | — |