BetterNormalCrafterWrapper
A ground-up ComfyUI inference implementation for NormalCrafter. It uses the released NormalCrafter UNet/VAE weights and the Stable Video Diffusion image encoder/scheduler.
Nodes (5)
Video surface normals that don't flicker, without the VRAM gymnastics
The Load node that keeps your model downloads in one place
The polite way to hand VRAM back after generating normals
Flicker-free normals from a clip, with the VRAM dials exposed
Two dropdowns, and the folder layout they actually want
Better NormalCrafter Wrapper
A small ComfyUI wrapper around NormalCrafter with two nodes, explicit VRAM controls, and staged GPU inference.
Nodes
NormalCrafter - Load
Selects two model sources through dropdowns:
- NormalCrafter model — UNet + VAE
- Base model — CLIP image encoder + feature extractor + scheduler
The wrapper registers and uses ComfyUI's model category normalcrafter. It does not rely on one hardcoded filesystem root. Model discovery goes through:
folder_paths.get_folder_paths("normalcrafter")
This means the normal ComfyUI location and any normalcrafter directories configured through extra_model_paths.yaml are searched together.
The standard location is registered as:
<ComfyUI models dir>/normalcrafter/
A corresponding extra path can be configured like this:
my_models:
base_path: D:/AI/
normalcrafter: models/normalcrafter/
Every compatible direct child folder found below any registered normalcrafter root appears in the dropdown.
A local NormalCrafter repo is detected by:
unet/
vae/
A local SVD base repo is detected by:
image_encoder/
feature_extractor/
scheduler/
The official Hugging Face sources remain available in the dropdown. If selected, only the required subfolders are downloaded. Existing copies are searched across all registered normalcrafter roots first; otherwise the first writable root in ComfyUI's path order is used.
NormalCrafter - Generate Normals
Inputs:
max_resolution— maximum long edge processed by the modelwindow_size— temporal UNet context size; official default is 14step_size— distance between temporal windows; official default is 10chunk_size— shared CLIP/VAE encode/decode batch size; lower values use less VRAM
step_size must be less than or equal to window_size.
Memory model
Inference is deliberately staged:
- CLIP encoder on GPU
- VAE encoder on GPU
- temporal UNet on GPU
- VAE decoder on GPU
Only the active heavy model is GPU-resident. Full-video CLIP embeddings and latent videos stay on CPU between stages. After every inference call all heavy modules are moved back to CPU automatically; there is no separate offload node.
window_size mainly controls temporal-UNet VRAM. chunk_size mainly controls CLIP/VAE VRAM. max_resolution affects all stages strongly.
Structure
nodes.py # ComfyUI surface: exactly two nodes
normalcrafter_clean/model.py # ComfyUI path discovery + loading + GPU residency
normalcrafter_clean/inference.py # preprocessing + four-stage inference
normalcrafter_clean/unet.py # NormalCrafter-specific SVD UNet forward