ComfyUI Custom Extensions
Browse 8,530 ComfyUI extensions and 65,510 custom nodes.
[a/ImageReward](https://github.com/THUDM/ImageReward): Human preference learning in text-to-image generation. This is a [a/paper](https://arxiv.org/abs/2304.05977) from NeurIPS 2023
A video editor for MiniMax H3 inside a single ComfyUI node. Keyframes with independent strength dials (guide the motion instead of pinning it), waypoints the clip passes through, and a fullscreen timeline editor with filmstrip cards, framing tool, cast members that persist across clips, and free Creative-Commons image/audio search built in. Continue a clip with real continuity: the previous render's tail frames AND audio are pinned at the new head, so motion and the actual waveform carry through the join — with an auto mode that chains clips hands-free. Finished clips land on a reel you can trim, crossfade, level, score with a soundtrack lane plus three fx lanes, and export as one video. Also includes v2v restyling, soft denoise zones (differential diffusion with per-frame masks), regional prompting and a temporal LoRA blend.
See what your seeds have in mind before you spend the steps: quick previews right on the node - click your favourite and only that image gets rendered. Exact continuation, no wasted steps, works with any model.
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion ,you can use a wrapper node it in comfyUI
OmniSVG: A Unified Scalable Vector Graphics Generation Model,you can try it in ComfyUI.
aesthetic for comfy ui
The CLIPSeg node generates a binary mask for a given input image and text prompt. NOTE:This custom node is a forked custom node with hotfixes applied from the [a/original repository](https://github.com/biegert/ComfyUI-CLIPSeg), which is no longer maintained.
ComfyUI nodes for running GGUF quantized Qwen2.5-VL models using llama.cpp
NODES: Auto-LLM-Text-Vision, Auto-LLM-Text, Auto-LLM-Vision
ComfyUI nodes for SkyReels-A2 model.
A NSFW/Safety Checker Node for ComfyUI
ComfyUI-Ranking精选社区,一个排行榜工具
A workflow subscription and version management hub for ComfyUI.
Custom Nodes for ComfyUI to Projects Texture into your 3D models.
A workflow browser & organizer for ComfyUI that operates on the real user/default/workflows folder plus any number of extra disk locations. Thumbnail gallery (auto-normalized 800x450, set/capture/drag-drop) and details List view with sortable columns including Tags. Tags: dedicated sidebar pane with live counts, chip-input editor with autocomplete, drag-and-drop assign, drag-and-drop reorder, custom/A-Z/Z-A sort, global rename/delete, per-workflow Copy/Paste/Clear. Search with focused (current view) default and a Global toggle to scan every root. Stackable Name/Date sort on cards. Favorites and per-workflow descriptions, full file/folder CRUD with multi-select (Ctrl/Shift range) and mandatory delete confirmation, and a Save button that recognizes the workflow you already have open. UI extension only — adds no graph nodes.
A collection of ComfyUI directory automation utility nodes. Directory Get-It-Right adds a GUI directory browser, and a smart directory loop/iteration node that supports regex + file extension filtering + sorting methods.
ComfyUI node to apply the ResAdapter Unet patch for SD1.5 models
Prompt Info
The World Weaver System: True AI character consistency using Textual Inheritance. Maintain unshakeable character identity (face, body, essence) across radical changes in pose, clothing, and scene without LoRAs, IP-Adapters, or ControlNet. This repo contains the Character Vault and Prompt Helper components.
Expose your workflows into HTTP endpoints directly from ComfyUI itself.
Experimental Windows bridge for NVIDIA DLSS Super Resolution, neural rendering, and Frame Generation in ComfyUI.
Connect to any Draw Things gRPC server
This is a ComfyUI plugin that provides a user interface of AudioMass, originally developed by [a/AudioMass](https://github.com/pkalogiros/audiomass)
Main nodes of the Level Pixel company (aka levelpixel, LP). Includes convenient nodes for working with images from folders; counting files in a folder; cleaning memory; tag filters. Model Unloader, LLM Unloader, Free memory, Tag Filters, Tag Category Filters, Tag Choice Parser, File counter, Image Loader From Path (with counters), Image Remove Background based on RemBG, Autotagger.
A ComfyUI custom node plugin for prompt gallery management, prompt selection, and image saving with categorization and cover image support.
A modern GUI-based color picker for ComfyUI nodes. Features visual spectrum, HEX/RGB inputs, eyedropper tool, and favorite colors support.
The ComfyUI Mask Bounding Box Plugin provides functionalities for selecting a specific size mask from an image. Can be combined with ClipSEG to replace any aspect of an SDXL image with an SD1.5 output.
A visual tool for prompt randomization and advanced combinations inside of your ComfyUI workflows.
ComfyUI GLM-4 Wrapper. This powerful tool enhances your prompt engineering process by allowing users to easily construct detailed, high-quality prompts for image/video generation based on user image and/or user prompts.
Registered outpainting nodes for the yijunwang2/krea2-outpaint LoRA on Krea 2 Turbo. Places the source reference into the target latent grid at a bbox (RoPE frame i+1) with isolated-ref KV cache, then restores exact pixels with a feathered seam. Reference-attention machinery adapted from ostris/ComfyUI-Krea2-Ostris-Edit (MIT); placement geometry from the outpaint release (Apache-2.0).
A ComfyUI custom node extension that integrates the Janus-Pro-7B vision-language model from DeepSeek AI on your's local computer, enabling powerful image understanding and multi-turn conversation capabilities.
Nodes:MultiLora Loader, Lora Text Extractor. Provides a node for assisting in loading loras through text.
using diffree: Text-Guided Shape Free Object Inpainting with Diffusion Model
LTX-2 is the first DiT-based audio-video foundation model that contains all core capabilities of modern video generation in one modelsynchronized audio and video
no desc
NODES:TTP_Hunyuan3DNode, TTP_SquareImage, TTP_GIFViewer
Adapt for Hunyuan now NOTE: The files in the repo are not organized, which may lead to update issues.
Convert Text-to-Speech inside ComfyUI using [a/Piper](https://github.com/rhasspy/piper)
Provides model-agnostic ComfyUI nodes for local repainting, detail repair, product/logo/text refinement, and paste-back compositing.
Semantic conditioning bridge for MiniMax H3.
TK Toolkit (formerly Anima Toolkit) — all-in-one Anima/SD toolkit for ComfyUI: LoRA management (batch apply + Civitai metadata/downloads), Danbooru search & waterfall gallery, bilingual prompt cards with Chinese tag autocomplete, an output gallery over ComfyUI/output, and batch workflow utilities.
NODES: AGSoft Empty Latent, AGSoft ImageRes, AGSoft ImageResMP, AGSoft ImagegPad for Outpainting, AGSoft ImagePad for OutpaintingAdv, AGSoft Image Crop, AGSoft Image Concatenate, AGSoft Image ConcatenateFromBatch, ...
Amazon Bedrock is a fully managed service that offers a choice of high-performing foundation models (FMs) from leading AI companies. This repo is the ComfyUI nodes for Bedrock service. You could invoke the foundation model in your ComfyUI pipeline.
A specialized node for ComfyUI that enable advanced motion and animation capabilities for image as guider for video processing In Hunyuan Video.
Advanced Model Manager for ComfyUI — browse, search and download models from HuggingFace and GitHub workflows directly from the ComfyUI interface.
Experimental and mathematically unsound (but fun!) sampling for ComfyUI. Feel free create a question in Discussions for usage help: OCS Q&A Discussion[w/Status: In flux, may be useful but likely to change/break workflows frequently. Mainly for advanced users.]
Legacy ComfyUI custom nodes for Boogu-Image; recommend using native ComfyUI support instead for new installations.
This repository is a quick port of [a/Resynthesizer](https://github.com/bootchk/resynthesizer) to ComfyUI. Resynthesizer is the open-source implementation of a texture generation technique proposed by Paul Harrison in 2005, especially useful for removing an object from an image (inpainting), which is most likely close to what Photoshop uses to for the content aware fill feature. Note that this is not using a diffusion model to inpaint, as opposed to many techniques of today, which makes it very fast and predictable, but sometimes yields worse results.
Fit, stylize, transfer, and camera-drive SAM3D Body MHR pose data in ComfyUI.
A powerful ComfyUI extension node that allows you to add various exquisite artistic text effects to your images, supporting a wide range of text styles and effects.
Node to enable Telegram in ComfyUI.
Minimal single-image MiniMax H3 photo editing with native or semantic references
Restore composition diversity for distilled diffusion models. Training-free frequency-domain phase injection.
A professional-grade toolkit providing lightweight APIs to integrate leading Large Language Models (LLMs) including DeepSeek, Qwen and GPT directly into ComfyUI.
GLM4 Vision Integration
Custom Graph Sigma is a ComfyUI custom node that provides an interactive spline-based curve editor for visually creating and exporting custom sigma schedules. This is especially useful for controlling the noise schedule or custom step values in diffusion models and other workflows that use a sequence of values over time or steps.
A custom node that makes prompt editing easier by allowing phrase switching with just mouse operations.
ComfyUI nodes for Qwen3-ASR (0.6B/1.7B) and ForcedAligner. Supports high-accuracy ASR and language identification for 52 languages/dialects, including 22 Chinese dialects and various English accents. Features word-level timestamps, long audio transcription, and VRAM-optimized inference.
Nodes:GradientPatchModelAddDownscale (Kohya Deep Shrink).
对 photodoodle 的 comfyui 原生实现
Nodes:BreakFrames, GetKeyFrames, MakeGrid.
ComfyUI-KepOpenAI is a user-friendly node that serves as an interface to the GPT-4 with Vision (GPT-4V) API. This integration facilitates the processing of images coupled with text prompts, leveraging the capabilities of the OpenAI API to generate text completions that are contextually relevant to the provided inputs.
A comprehensive suite of custom nodes for building structured JSON prompts for FLUX.2 image generation with precision and control.
Nodes:OpenAI DALLe3, OpenAI Translate to English, String Function, Seed Generator
Nodes:Noodle webcam is a node that records frames and send them to your favourite node.
Nodes:LatentGarbageCollector. This ComfyUI custom node flushes the GPU cache and empty cuda interprocess memory. It's helpfull for low memory environment such as the free Google Colab, especially when the workflow VAE decode latents of the size above 1500x1500.
CLIPTextEncodeBLIP: This custom node provides a CLIP Encoder that is capable of receiving images as input.
Four nodes for MiniMax H3: type a plain prompt, name your pictures, clips and sounds with @, lock exact dialogue, and get the structured brief H3 wants in its own published format, checked before it runs. Nothing to start beside ComfyUI. Needs any OpenAI-compatible model you run or have access to.
ComfyUI custom nodes for OpenMOSS MOSS-TTS Local Transformer v1.5 with model catalog downloads, Whisper prefix transcription, continuation, voice cloning, and AIMDO memory tracking
This is a ComfyUI custom node used to convert Qwen-Image LoRA files trained on the ModelScope platform to a format that ComfyUI can recognize.
A powerful ComfyUI node for rendering text with advanced styling options, including full support for Persian/Farsi and Arabic scripts.
ComfyUI_Gemini_Flash is a custom node for ComfyUI, integrating the capabilities of the Gemini 1.5 Flash model. This node supports text and vision-based prompts, allowing users to analyze and adapt images to text prompts for text2image tasks.
Physics-based post-processing using depth maps to simulate atmospheric perspective, light transport, and lens imperfections for photorealistic results.
VoxCPM锛歍okenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning锛寉ou can use this node,easy infer and easy train
Spectrum: training-free diffusion sampling acceleration via Chebyshev polynomial feature forecasting. Drop-in KSampler replacement that skips transformer blocks on predicted steps for ~2-3x speedup.
A ComfyUI extension for Qwen-VL series large language models, supporting multi-modal functions such as text generation, image understanding, and video analysis.Support for Qwen2-VL, Qwen2.5-VL.
Abstract Syntax Trees Evaluated Restricted Run (ASTERR) is a Python Script executor for ComfyUI. [w/Warning:ASTERR runs Python Code from a Web Interface! It is highly recommended to run this in a closed-off environment, as it could have potential security risks.]
用于处理长文本文件并生成语音的ComfyUI插件,支持音频分离、文本分割、音频拼接等功能
Nodes: SDXLResolutionPresets. Easy access to the officially supported resolutions, in both horizontal and vertical formats: 1024x1024, 1152x896, 1216x832, 1344x768, 1536x640
ComfyUI-Bagel is now available in ComfyUI, BAGEL is an open‑source multimodal foundation model with 7B active parameters (14B total) trained on large‑scale interleaved multimodal data. [w/Don't install together with neverbiasu/ComfyUI-BAGEL simultaneously.]
ComfyUI node for loading a MiniMax H3 FL2VA/Ref2VA hybrid model with ComfyUI DynamicVRAM/AIMDO-aware loading and matched checkpoint memory behavior.
A small node pack containing various things I felt like ought to be in base comfy-UI. Currently includes Some image handling nodes to help with inpainting, a version of KSampler (advanced) that allows for denoise, and a node that can swap it's inputs. Remember to make an issue if you experience any bugs or errors!
Easy ComfyUI custom nodes for Meta Sapiens2 segmentation, normal, pointmap, and pose workflows.
Multi-GPU Ulysses sequence-parallel inference for MiniMax-H3 video and audio generation across 2/4/7/8 GPUs. Requires ComfyUI commit c194dd00 or newer and Linux or WSL2 with NCCL; native Windows is not supported.
Architecture-aware LoRA loader for FLUX.2 Klein in ComfyUI with automatic per-layer strength calibration.
A text-to-speech plugin used under ComfyUI. It utilizes the Microsoft Speech TTS interface to convert text content into MP3 format audio files.
Choose which image to use based on the keywords in the prompt.
Custom nodes for deforum workflows
A collection of custom nodes for ComfyUI.
A simply node for hooking in to openAI API based servers via comfyUI
Conveniently control parts of text prompts with custom UI. Pack includes loaders from txt and csv files, dynamic text concatenation tool and easy-to-use input node
A ComfyUI custom node enabling **Flash Attention 1** on legacy NVIDIA GPUs (Tesla V100, T4) that lack Compute Capability 8.0+ required by FlashAttention-2.
Utility custom nodes for special effects, image manipulation and quality of life tools.
This extension integrates interactive multi-layer canvas editing, multimodal LLM intelligent processing (supporting mainstream models like DeepSeek, Qwen, Kimi, Wan), professional-grade image processing toolchains, video generation orchestration systems (supporting Wansiang series reference-based video, image-to-video, and keyframe-to-video generation), and visual data tools, providing end-to-end support for AI image and video generation workflows from creative conception to fine editing. With advanced PSD import, BiRefNet intelligent matting, real-time layer transformations, unified configuration-based video generation orchestration, and secure API key management, it significantly enhances creative efficiency and output quality. Additionally, it offers workflow customization features like node color management and label nodes to improve visual organization and personalization of complex workflows.
Nodes for auto download models from Hugging Face using their filenames as part of workflows
ACE-Step 1.5 music generation nodes for ComfyUI - Generate high-quality music from text
ACE-Step 1.5 music generation nodes for ComfyUI - Generate high-quality music from text
Powerful Flux-Kontext image generation custom node for ComfyUI, using the official RabbitAI API. Supports text-to-image, image-to-image, and multi-image-to-image generation. Supports concurrent generation.