ComfyUI Extension: ComfyUI-TRELLIS-HiCache

Authored by Archerkattri

Created

Updated

2 stars

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Training-free TRELLIS image-to-3D acceleration: forecast the flow-matching velocity on skipped DiT steps (HiCache Hermite / HiCache++ DMD, via hicache-pp) across both the sparse-structure and SLaT stages. ~2x faster, near-lossless. Drop between the TRELLIS loader and sampler.

Looking for a different extension?

Custom Nodes (1)

README

ComfyUI-TRELLIS-HiCache

Training-free acceleration for TRELLIS image-to-3D in ComfyUI. It forecasts the flow-matching velocity on skipped DiT steps instead of running the transformer, on both TRELLIS stages (sparse-structure and SLaT), via the hicache-pp library.

Pairs with smthemex/ComfyUI_TRELLIS (the MODEL_TRELLIS pipeline type). Same idea as ComfyUI-HiCache for Hunyuan3D.

What it does

TRELLIS samples each stage with a flow-Euler loop that calls a DiT once (or twice, under classifier-free guidance) per step. TRELLIS HiCache Accelerate replaces the two flow DiTs with a wrapper that runs the transformer on a schedule and forecasts the velocity on the steps in between:

  • hermite — HiCache (dual-scaled physicist's Hermite polynomial, arXiv:2508.16984).
  • dmd — HiCache++ (Dynamic Mode Decomposition / Prony exponential basis).
  • auto — holdout-selected per step.

Two TRELLIS-specific details are handled correctly: the timestep schedule runs 1 -> 0 (so run boundaries are detected by direction reversal, not a fixed threshold), and classifier-free guidance issues the conditional and unconditional forwards separately inside a guidance interval, so the patch keeps two parallel forecast states and routes each forward to the right one. The SLaT stage returns a sparse tensor whose active-voxel layout is fixed during a run, so the forecast runs on its .feats and the sparse tensor is rebuilt from the last computed step.

Measured (RTX 5090, TRELLIS-image-large, demo image, 25+25 steps)

| config | speedup | Chamfer vs stock (unit-cube units) | |---|---|---| | interval=2, both stages | 2.1x | 0.0059 (near-lossless) | | interval=3, both stages | 2.5x | 0.0145 (more aggressive) |

interval=2 is the default. Chamfer is the symmetric mean nearest-neighbour distance between the stock and accelerated Gaussian point clouds; the object spans ~1.0, so 0.0059 is ~0.6% of its extent. The gaussian count shifts more than the surface does, because TRELLIS' active-voxel threshold is sensitive near the boundary — surface fidelity is the metric that matters and is what the Chamfer column reports.

Install

In ComfyUI: install via the ComfyUI Manager (search "TRELLIS HiCache"), or:

cd ComfyUI/custom_nodes
git clone https://github.com/Archerkattri/ComfyUI-TRELLIS-HiCache
pip install hicache-pp

Use

Trellis_LoadModel -> TRELLIS HiCache Accelerate -> Trellis_Sampler

Set enabled = Off to bypass and restore the stock DiTs. The node never mutates the pipeline it is given (copy-on-patch), so a cached node output always owns its own configuration.

Validation

tests/test_patch.py unit-tests the patch logic with a dummy DiT (no ComfyUI, no GPU). tests/validate_gpu.py is the end-to-end GPU check that produced the table above (loads a real TRELLIS pipeline, applies the patch, compares geometry and wall-clock against stock).

Apache-2.0.

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Learn more