Extensions/BetterNormalCrafterWrapper
ComfyUI Extension

BetterNormalCrafterWrapper

A ground-up ComfyUI inference implementation for NormalCrafter. It uses the released NormalCrafter UNet/VAE weights and the Stable Video Diffusion image encoder/scheduler.

By kaski23·Created 3 months ago·Updated 10 days ago· 0
kaski23/BetterNormalCrafterWrapper
Nodes5
On cloudLocal install
CategoryNormalCrafter/Clean, NormalCrafter
Stars0
Updated10 days ago
Readme

Better NormalCrafter Wrapper

A small ComfyUI wrapper around NormalCrafter with two nodes, explicit VRAM controls, and staged GPU inference.

Nodes

NormalCrafter - Load

Selects two model sources through dropdowns:

  • NormalCrafter model — UNet + VAE
  • Base model — CLIP image encoder + feature extractor + scheduler

The wrapper registers and uses ComfyUI's model category normalcrafter. It does not rely on one hardcoded filesystem root. Model discovery goes through:

folder_paths.get_folder_paths("normalcrafter")

This means the normal ComfyUI location and any normalcrafter directories configured through extra_model_paths.yaml are searched together.

The standard location is registered as:

<ComfyUI models dir>/normalcrafter/

A corresponding extra path can be configured like this:

my_models:
  base_path: D:/AI/
  normalcrafter: models/normalcrafter/

Every compatible direct child folder found below any registered normalcrafter root appears in the dropdown.

A local NormalCrafter repo is detected by:

unet/
vae/

A local SVD base repo is detected by:

image_encoder/
feature_extractor/
scheduler/

The official Hugging Face sources remain available in the dropdown. If selected, only the required subfolders are downloaded. Existing copies are searched across all registered normalcrafter roots first; otherwise the first writable root in ComfyUI's path order is used.

NormalCrafter - Generate Normals

Inputs:

  • max_resolution — maximum long edge processed by the model
  • window_size — temporal UNet context size; official default is 14
  • step_size — distance between temporal windows; official default is 10
  • chunk_size — shared CLIP/VAE encode/decode batch size; lower values use less VRAM

step_size must be less than or equal to window_size.

Memory model

Inference is deliberately staged:

  1. CLIP encoder on GPU
  2. VAE encoder on GPU
  3. temporal UNet on GPU
  4. VAE decoder on GPU

Only the active heavy model is GPU-resident. Full-video CLIP embeddings and latent videos stay on CPU between stages. After every inference call all heavy modules are moved back to CPU automatically; there is no separate offload node.

window_size mainly controls temporal-UNet VRAM. chunk_size mainly controls CLIP/VAE VRAM. max_resolution affects all stages strongly.

Structure

nodes.py                         # ComfyUI surface: exactly two nodes
normalcrafter_clean/model.py     # ComfyUI path discovery + loading + GPU residency
normalcrafter_clean/inference.py # preprocessing + four-stage inference
normalcrafter_clean/unet.py      # NormalCrafter-specific SVD UNet forward