ComfyUI-Transformers
ComfyUI-Transformers is a cutting-edge project combining the power of computer vision and natural language processing to create intuitive and user-friendly interfaces. Our goal is to make technology more accessible and engaging.
Nodes (33)
What's That Sound? Audio Classification in ComfyUI
Whisper ASR
Chat With a Model Inside ComfyUI (Conversational)
Depth Estimation
Document QA
Feature Extraction
Fill Mask in ComfyUI
Convert Numbers to Text in ComfyUI (Float to String)
What Is This Image? Image Classification in ComfyUI
Image Feature Extraction
Segmentation in ComfyUI
Image-Text to Text in ComfyUI
Upscale or Transform Images With One Node (Image to Image)
Auto-Caption Any Image Inside ComfyUI (Image to Text)
Turn a Number Into Text in ComfyUI (Int to String)
The Load Node for ComfyUI's Depth Workflow
Automatic Mask Generation
Drop a DETR model into your graph
Extract Answers From Text in ComfyUI (Question Answering)
Score how close two texts are — the node that can gate your workflow
The boring utility that keeps your graph from screaming at you
StringToInt — convert a text value into an integer ComfyUI can use
Ask your data a question in the middle of a diffusion workflow
Give your graph a sentiment check before it spends a generation
The prompt-generator node — and the trap that makes it lie to you
Text to Speech with Bark
Pull the people, places, and organizations out of any text
What's happening in that clip? The node that labels video content
Ask your generated image a question and let it answer
Classify any audio into labels you invent on the spot
Sort text into categories you invent — no training required
Does this image contain a cat? Ask CLIP, no training involved
Zero-Shot Detection With OWL-ViT
🛠️ Installation
cd custom/nodes
git clone https://github.com/kadirnar/ComfyUI-Transformers