comfyui-conduit-optimizer
Non-linear inference optimization for ComfyUI: 4-tier VRAM, speculative generation, precision routing
Nodes (13)
Where your optimizations are supposed to happen
Let VRAM pick your batch size
The reset button for the Conduit caches
The mode selector at the head of the Conduit optimizer
Stop re-encoding the same prompt
Choose your precision without melting your GPU
Identical runs, instant results
Speculative generation, on paper
The VRAM temperature scale that mostly describes itself
Half-precision attention without touching your checkpoint
The deterministic cache you configure but can't inspect
The node that reads your prompt and guesses what you're doing
Roll several seeds, keep the best
CONDUIT - ComfyUI Inference Optimizer
"Optimize the flow of your inference"
A production-ready optimization system for ComfyUI that implements non-linear, non-binary optimization strategies for faster generation with lower VRAM usage.
Nodes
| Node | Purpose | Key Feature | |------|---------|-------------| | ConduitCore | Workflow Optimizer | DAG analysis, async scheduling | | ConduitPool | VRAM Manager | 4-tier memory temperature gradient | | ConduitGate | Precision Router | Dynamic FP32/FP16/FP8 routing | | ConduitPath | Speculative Generator | Multi-branch generation with pruning | | ConduitSeal | Deterministic Cache | Checksum-verified result caching | | ConduitSense | Type Detector | Auto-detect workflow type | | ConduitApply | Apply Optimizations | Combine configs and execute |
Quick Start
- Add ConduitCore to set optimization mode (balanced/speed/quality/memory)
- Add ConduitGate to configure precision routing
- Connect to ConduitApply before your model
- Run workflow
Optimization Modes
Speed Mode
- Aggressive FP8 precision on RTX 40 series
- Async model preloading
- Estimated 2-3x speedup
Quality Mode
- Full FP32 precision for attention
- No early exit from sampling
- Best visual quality
Memory Mode
- Aggressive VRAM offloading
- Streaming decode
- Tile processing for large images
Balanced Mode (Default)
- Adaptive precision based on operation type
- Smart preloading without VRAM pressure
- Good balance of speed and quality
VRAM Temperature System (ConduitPool)
HOT (GPU VRAM) - Active inference
WARM (Pinned) - Next 2-3 models, instant transfer
COLD (RAM) - Recently used, ~500ms load
ARCHIVE (Disk) - Rarely used, predictive prefetch
Speculative Generation (ConduitPath)
Generate N branches, score at checkpoints, prune losers:
4 branches @ 25% → Score → Keep 2
2 branches @ 50% → Score → Keep 1
1 branch @ 100% → Output
Result: High-quality output at ~50% compute cost
Requirements
- ComfyUI (latest)
- PyTorch 2.0+
- CUDA 11.8+ (for FP8 on RTX 40 series)
Hardware Optimization
Automatically detects and uses:
- RTX 40 series FP8 TensorCores (1320 TFLOPS)
- BF16 on Ampere+ GPUs
- Mixed precision for optimal performance
License
Apache 2.0 - See LICENSE file
Built with advanced optimization research