ComfyUI Node
HunyuanVideo Vision Encode
A ComfyUI node in HunyuanVideoWrapper1.5 with 8 inputs and 1 output.
HunyuanVideo Vision Encode
- vision_encoder
- hyvid_cfg
- latents_dict
- reference_image
- vision_states
◄target_dtypebfloat16►
◄enable_offloadingtrue►
◄vision_num_semantic_tokens729►
◄vision_states_dim1152►
CategoryHunyuanVideoWrapper1.5
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| vision_encoder | HYVID15VISIONENCODER | — | |
| hyvid_cfg | HYVID15CFG | — | |
| latents_dict | HYVID15LATENTSDICT | — | |
| target_dtype | COMBO | bfloat16 | 9 options: float32, float64, float16, bfloat16, uint8, int8, +3 |
| enable_offloadingopt | BOOLEAN | true | — |
| reference_imageopt | IMAGE | — | |
| vision_num_semantic_tokensopt | INT | 729 | — |
| vision_states_dimopt | INT | 1152 | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| vision_states | HYVID15VISIONSTATES | — |