XB_ToolBox
A comprehensive ComfyUI extension suite encompassing both front-end interaction and low-level memory scheduling, designed to help beginners master workflows and simplify local deployment.
Nodes (154)
Slice audio to exact frames for video models — no manual math
The audio slicer with a waveform you can actually click on
Stitch two voice clips into one track, gap and all
The audio slicer for when two voices actually talk over each other
Erase AI gibberish in comic bubbles and typeset real dialogue
Load one image at a time from a folder, stepper built in
Merge loose images into one batch tensor, mute-proof
Feed source video and reference images into Bernini without OOMing
One dropdown replaces a 30-node Bernini prompt-switching subgraph
A sticky note for your ComfyUI graph that can't break anything
Push transformer blocks to RAM so big video models fit on your GPU
See your spatial and temporal chunks before you blow your VRAM budget
INT8 text encoders for AMD cards that were dying on VRAM
Turn a CLIP model name into a wire so other nodes can read it
Strip dialogue out of your comic prompt before the model sees it
Place comic text at exact coordinates — narration, SFX, anything
Trim a reference clip by wall-clock time, not samples
Clone a voice once, then make it speak any of ten languages
Two cloned voices, one script, one mixed conversation track
Tell the cloned voice how to sound, in plain words
The CosyVoice3 loader that actually gets the model into ComfyUI
Turn a voice clip into a reusable .pt speaker file, once
Synthesize in a saved voice, no reference audio required
Same cloned voice, but now tell it how to say the line
Respeak someone else's words in your own voice
Clone a voice from one audio clip, no training, no transcript required
A clean control panel for your messy workflow, without rebuilding it
Two voices, two audio tracks, one synchronized frame count
Frame math and audio slicing for a one-person talking head, in one node
Load two INT8 text encoders on AMD — the memory trick Flux users want
A 20-lane wire bundle for when your graph looks like spilled spaghetti
Stuff up to nine reference images into a Flux conditioning, in one node
Hailuo H3 resolution and frame math, without the mental arithmetic
Cut a person out of any image with a GPU-accelerated mask
Load the ONNX segmentation model that the segmenter needs
One hub for image size, batch, and strength — with dims that can't go invalid
Scale up to nine images at once with one shared setting
Save your INT8 text encoder once, reuse it forever
Stack multiple LoRAs on an INT8 model without the slot limit
A LoRA loader that respects your INT8 model
Save the INT8 model so you never re-quantize it again
Pre-load a LoRA so it gets baked into the INT8 weights
One dropdown instead of ten toggles for your K2 style switch
The stock KSampler, plus a before-sampling VRAM cleanup knob
KSamplerAdvanced, with a pre-run VRAM cleanup bolted on
Split one list of strings into N separate text wires
Turn LLM-detected boxes into a soft, feathered mask
LLM-detected boxes, fed straight into the Impact Pack detailer
Pluck one box out of a list of LLM-detected boxes
Reset the LLM's memory so it stops 'remembering' old runs
Run your local GGUF model as a full instruct + vision node
Turn your LLM's bounding-box JSON into real BBOX wires
A complete MiniMax H3 prompt builder in one dropdown
The multi-reference MiniMax H3 prompt preset
The GGUF model loader that powers the whole XB-llama stack
The sampling knobs for your local LLM, tuned for storyboards
Pull one value out of your LLM's JSON without a Python node
Model-specific system prompts, one dropdown at a time
Turn a story idea into a frame-by-frame storyboard prompt
Story idea in, frame prompts out
Draw cards from your storyboard, one shot at a time
Same storyboard gacha, plus it sniffs your characters for you
Kick the LLM off your GPU when you're done with it
Strip the ``` fences off your LLM's JSON
Unlimited-length LTX 2.3 lipsync, stitched from segments
One value box that snap-clamps itself to the mode you're in
The VRAM radar that parks itself in the corner of your screen
Model, CLIP, eight LoRA slots, VAE, Sage, block swap — in one node
V1's loader, but the model dropdown speaks GGUF
Two models, one loader — high-noise and low-noise in a single node
High/low dual-model loader, GGUF edition
Dual CLIP and dual VAE in one node
LTX 2.3's dual-CLIP, dual-VAE stack, quantized to fit
Four images in, one frame sequence out — MSR for short dramas
Mute or bypass whole nodes from a boolean — without right-clicking
One output, many optional upstreams — pick the live one by name
KSampler, with a memory-cleanup button welded to the front
KSamplerAdvanced with the pre-run memory sweep
Tiled, temporal VAE decode for video latents that won't fit
The manual 'clear VRAM' button you can drop anywhere in a graph
XB_ROCmSamplerCustom is just ComfyUI's SamplerCustom with a cleanup switch
The same old advanced sampler, now with a panic button for your VRAM
The 'ROCm' VAE decoder that secretly tiles everything
Tiling your way out of video-VAE OOM, frame by frame
Get your reference images into latent space without the OOM cliff
Turn one storyboard script into per-scene prompts plus matching characters
Free attention speedup, with a preset dropdown for your exact GPU
The 'golden duo' that runs big models on small cards
Seeing your video's memory footprint before it OOMs
A sampler that does exactly what it says, plus a VRAM shovel
ComfyUI's most flexible sampler, gift-wrapped with a cleanup dropdown
One click from storyboard grid to individual scene frames
Join N strings without a tangle of text nodes
Tell a picture what to change, no mask required
Multi-image Qwen editing, with the vision-language plumbing done for you
Trading speed for VRAM, one transformer block at a time
INT8 diffusion models for cards that can't fit the real thing
One dropdown, one wire, every model name you need
The plainest node in the pack, and that's the point
The decoder you actually want on a 16GB card
Decode huge latents without OOM — this is ComfyUI's tiled decoder with a cleanup switch
Stock VAE encode, with a VRAM cleanup knob you can actually see
Inpainting encode with a cleanup knob — same masked latent, less OOM
Tiled VAE encode for big images and video — with the cleanup switch ComfyUI forgot
The VHS video saver that got ported into XB_ToolBox — with GIF/WebP shortcuts
Load any video into frames with VHS compatibility — and the preview bug is actually fixed
Stitch up to ten video segments into one long clip — on the CPU, so it won't OOM
One hub for width, height, frames and FPS — with the 1+8N and golden-bucket rules built in
Does your card actually fit that 22B model? This node does the math so you don't have to
Wan 2.2 reference + control-video conditioning, prebuilt so you don't hand-build concat latents
A start image in, a Wan 2.2 latent out — the I2V setup that doesn't fight the VAE
One node to configure a whole motion-transfer run — the Animate parameter bus
Chaining motion transfer past the 81-frame wall — the Animate relay node
The better relay — one node that loops as many Animate segments as you need
The engine under Wan Animate motion transfer — reference image + pose video in, sampled video out
Swap transformer blocks to RAM so a 14B Wan model fits your card
Wan camera-control conditioning, with a start frame and tiled VAE encode
Torch.compile for Wan, as a settings bundle — and the reason AMD users should skip it
One node that turns a music file into dance prompts and audio features for Wan Dancer
Dance conditioning for Wan-Dancer — the official logic, tiled so it fits in VRAM
The Wan Dancer switcher
The tile-stride dial that saves your run
Wan first/last frame animation that stays on rails
One node for control-video conditioning
Bookend-pinned editing for Wan 2.2
Audio-driven human motion with Wan HuMo
The node that turns your start image into Wan's I2V conditioning
The relay node that chains Wan clips into a continuous video stream
Relay_count turns N nodes into one
One bus for the whole talking-head pipeline
A segment generator that knows where the audio ends
Three ways to pick a segment's first frame
Per-segment reference images from your input folder
Relay_count for talk-till-the-audio-ends
One node for InfiniteTalk audio-to-video segments, single or dual speaker
Dual-speaker InfiniteTalk chunks with per-person masks
The no-surprises InfiniteTalk engine
The Wan loader that makes fp8, SageAttention and block-swap choices legible
The Wan Param Bus that ends noodle-wire chaos
Multi-reference conditioning with Phantom Subject
Stitch a whole video from first-and-last frames
The Wan sampler that cleans up after itself
One control panel for a whole SCAIL relay chain
Turn SCAIL into a five-minute video
Plain SCAIL, one segment, no scaffolding
SCAIL-2 with the relay bookkeeping built in
Make Wan video follow the audio
Keep the audio going, extend the video
The Wan T5 loader that reads umt5-xxl, not CLIP
Two prompt boxes, one T5 pass
Region-aware Wan video editing
Decode long Wan latents without the OOM
The Wan VAE loader with the free VRAM trim
A wire you can attach to see what your data weighs
🧰 ComfyUI XB-BOX (XB_ToolBox)
🌟 Core Philosophy
The primary goal of XB_ToolBox is to help AI beginners new to ComfyUI quickly master workflows, making local deployment and execution simpler and more convenient. It is a comprehensive ComfyUI extension suite encompassing both front-end interaction and low-level memory scheduling.
- Unified Parameters: Integrates frequently used but scattered parameter nodes, allowing you to set essential image and video parameters from a single hub.
- Visual Learning: Provides visual chunking and parameter preview nodes, enabling beginners to intuitively understand the concepts of VAE/Encoder chunks, temporal slicing, and overlap effects.
- UX Enhancements: Introduces a user-friendly "Mirror Clone Console" (Dashboard) that clones major operation nodes from complex workflows into a unified, clean, and highly efficient control panel.
- Extreme VRAM Optimization: Through "Spatiotemporal Chunk Hijacking", "Dynamic VRAM Offloading for UNet/Checkpoints", and "Deep LiteGraph Customization", it allows consumer-grade GPUs to smoothly run massive 14B~22B 3D-DiT video models that would normally cause instant OOM (Out of Memory).


✨ Core Nodes
1. 🎬 Media Parameters & Spatiotemporal Visualization
- 🖼️ Image Parameters Master: Controls core image generation parameters (Width, Height, Batch Size, Strength) in one place. Enforces safe steps to prevent invalid inputs and automatically aligns dimensions when a fixed aspect ratio is selected.
- 🎬 Video Parameters Master: Manages Width, Height, Frames, and FPS (Integer/Float). Features a geek-level "Auto-Manual Transmission Engine" that smoothly snaps to official pre-trained "Golden Buckets" (e.g., 480x832, 544x960) and strictly locks the physical frame count to the safe
1+8Nrule. Dynamically displays video duration based on frame rate. - 🧊 Spatiotemporal Chunk Visualization: An exclusive dual-zone radar! The 2D left panel visualizes spatial image chunks and overlap values, while the 3D right panel uses a painter's algorithm to render temporal cylinder stacks. Saves beginners from the exhaustion of blind parameter tuning.
- 🧊 Storyboard Image Slicer: Offers multiple modes to slice common 4-grid, 6-grid, 9-grid... images all at once and output multiple individual images. This makes image-to-video or start/end frame video generation much more convenient, and is highly effective for creating short dramas.
- 📟 VRAM Calculator & Data Radar: Predicts VRAM footprint for WAN/LTX models under different quantizations (FP8/GGUF) and weighs tensor volumes with MB-level precision. Significantly reduces trial-and-error costs by helping users choose models fitting their VRAM capacity.
2. 🧊 VRAM Optimization
- 🧊 Sampler Chunk Master: Tailored for WAN/LTX. Performs spatiotemporal dual-axis slicing (Spatial Tiles & Temporal Chunks) in the Latent space. Hijacks model parameters in-place with an built-in
rocm_optimizedextreme memory reclamation strategy, breaking the large-tensor OOM curse. - ✂️ Model Block Swap: Supports both UNet and Checkpoint modes. Forcibly unloads the first N Transformer Blocks and Text/Image embedding layers to system RAM, trading time for space.
- 🧹 VRAM Cleaner: Essential for pre-generation clearing. Delivers "nuclear-level cleanup" by deeply invoking
gc.collect()andtorch.cuda.empty_cache()to crush PyTorch cache fragmentation.
3. 🎛️ Workflow Deployment Assistants
- 🪄 XB Dashboard Zen: A "Mirror Clone Console" developed using low-level LiteGraph features. Batch-clones any scattered components in your workflow into a fixed panel. Features two-way data synchronization and direct arterial connections to create a clean, customized monitoring dashboard.
- 🎛️ Dynamic Bus: An ultra-compact N-in/N-out
AnyTypeuniversal bus node. Provides custom UI type labels and one-click channel addition/removal to save your workflow from "noodle wire" chaos.
📦 Installation
Method 1: ComfyUI Manager (Recommended)
Search for XB_ToolBox in the ComfyUI Manager and click install.
Method 2: Manual Git Clone
Navigate to your ComfyUI custom_nodes directory and run:
git clone https://github.com/wjluoxiao/XB_ToolBox.git
(Note: This extension relies purely on the native ComfyUI ecosystem through elegant Python/JS architecture. NO extra pip dependencies required! Plug and play.)
📝 License This project is open-sourced under the Apache-2.0 License. It guarantees the freedom to share while protecting the patent defense rights of the core architecture code.