HSWQ ConvRot INT8/ConvRot NVFP4 UNet Loader (DisTorch2)
When the quantized UNet still won't fit
- MODEL
First, the honest framing
HSWQ (Hybrid-Sensitivity-Weighted-Quantization) has roughly zero search traffic, and there's a reason: it's a single-author quantization project, not a format the community voted for. Searching the reddit corpus for "HSWQ" returns nothing at all, while the thing it's built on - DisTorch2, from ComfyUI-MultiGPU - shows up in hundreds of threads since early 2025.
You're here one of two ways: you downloaded an HSWQ pack from ussoewwin's HuggingFace repos and the stock UNet loader choked, or you're running a quantized model that fits your budget, not your VRAM.
What it actually does
Two jobs stapled together. First, it's the pack's HSWQ-aware UNet loader - it reads comfy_quant metadata, ConvRot (pre-rotated INT8/NVFP4) weights, and Krea2's txtfusion.projector marker, none of which stock ComfyUI loaders interpret. Second, it wraps that loader in the DisTorch2 backend ported from pollockjj's ComfyUI-MultiGPU: declare a virtual VRAM budget plus a donor device and DisTorch places blocks accordingly, so the whole UNet can live outside VRAM. It patches mm.load_models_gpu and ModelPatcher.partially_load, so offloaded blocks sit in host memory and only what's needed stays resident.
Upstream's DisTorch2 loaders use the stock UNETLoader as their base, so HSWQ metadata goes uninterpreted - hence the in-tree port (distorch_2.py, GPL-3.0).
Where this sits in the fit-vs-quality ladder: fp8 halves VRAM and is the default advice for 40-series and up, NVFP4 is Blackwell-only and hardware-accelerated, and INT8-ConvRot works on 20/30/40/50-series - which is why ComfyUI adopted it natively in v0.27.0. Use int8_tensorwise for Krea2 ConvRot INT8 packs, the Z Image or Krea2 ConvRot NVFP4 options for NVFP4 ones, and default to auto-detect.
The inputs you'll actually touch
unet_name (your safetensors under models/diffusion_models), weight_dtype, and hswq_bake.
hswq_bake is the one that trips people up. Per the author's tooltip: ON gives you the HSWQ path (HSWQ LoRA bake plus the legacy patcher), required for HSWQ-only formats like Hybrid ConvRot NVFP4. OFF keeps the stock ComfyUI path under DynamicVRAM, skips every HSWQ patch, and drops offloaded weights into shared VRAM. On Krea2 ConvRot INT8, OFF is the faster side - 5.34–5.51 s/it in the README. Mechanically, ON is what lets DisTorch placement run: the loader re-clones the patcher with disable_dynamic=True, because a dynamic patcher sails past the partially_load hooks DisTorch patches into.
The offload knobs: virtual_vram_gb is the budget carved off into the donor device and must be at least the size of the UNet - 14 GB or more for Krea2 ConvRot INT8, which is also why the README's floor is 64 GB of system RAM. donor_device is where that memory lives (leave it on cpu), compute_device picks the GPU doing the math, expert_mode_allocations lets you write the allocation by hand in compute;gb;donor form, and eject_models (on by default) evicts already-loaded models before this load - handy on heavy graphs.
Output is a single MODEL: into LoRA loaders, the pack's own Torch Compile node, then a sampler.
Install
ComfyUI Manager: search ComfyUI-HSWQ-Loader-and-Tools. Or by hand:
cd ComfyUI/custom_nodes && git clone https://github.com/ussoewwin/ComfyUI-HSWQ-Loader-and-Tools
Restart ComfyUI. The pack has a real requirements.txt (diffusers, transformers, peft, accelerate, plus a detection stack: insightface, facexlib, opencv-python, onnxruntime) and an install.py that pip-installs it after upgrading setuptools/wheel - that last part exists because older setuptools breaks transitive source builds on Python 3.12.
No model files ship; pull HSWQ weights from the author's HuggingFace repos into models/diffusion_models.
Where people get burned
Only Krea2 ConvRot INT8 is verified with DisTorch2. The README and changelog both say so; the other weight_dtype choices aren't tested here yet. If something misbehaves, that's why.
virtual_vram_gb too small is the classic failure. The backend logs the minimum it needs ("To prevent an OOM error, set 'virtual_vram_gb' to at least N"); ignore it and you get the OOM. Start at the UNet's size, not the default 4.
On a light workload this is slower than the plain HSWQ UNet loader - offload overhead is real, and the author says so. It wins when the job is heavy: high resolution, LoRAs, ControlNet, multi-model graphs. The README measured a 16 GB card (RTX 5060 Ti) doing a simple generation faster on the plain loader.
No SageAttention2 here. The DisTorch2 node drops the attention_accel selector its sibling has; for SA2 use the plain UNet loader, and don't stack it with Patch Sage Attention DM.
Second-generation failures are a known bug family in this pack. HSWQ's residual GPU and host memory isn't fully released by ComfyUI's generic unload, so the README requires General Purge VRAM V2 (from ussoewwin's ComfyUI-DistorchMemoryManager) at the end of the workflow with its HSWQ toggle on. Skip it and your second run breaks.
If you came from upstream MultiGPU, a wall of Pin error. lines is expected: cudaHostRegister returns "already registered" (rc 712) for weights the aimdo host buffers pinned. This pack handles that case and logs once.
Inputs (8)
| Name | Type | Default | Description |
|---|---|---|---|
| unet_name | COMBO | 0 options: | |
| weight_dtype | COMBO | 7 options: default, fp8_e4m3fn, fp8_e4m3fn_fast, fp8_e5m2, int8_tensorwise, Z Image ConvRot NVFP4, +1 | |
| hswq_bake | BOOLEAN | true | ON = HSWQ path (HSWQ LoRA bake + legacy patcher; required for HSWQ-only formats such as Hybrid ConvRot NVFP4). OFF = stock ComfyUI path under DynamicVRAM (faster; offloaded weights land in shared VRAM). |
| compute_deviceopt | COMBO | cpu | 1 options: cpu |
| virtual_vram_gbopt | FLOAT | 4.00–128 | — |
| donor_deviceopt | COMBO | cpu | 1 options: cpu |
| expert_mode_allocationsopt | STRING | — | |
| eject_modelsopt | BOOLEAN | true | — |
Outputs (1)
| Name | Type | Description |
|---|---|---|
| MODEL | MODEL | — |