ComfyUI Extension: ComfyUI-DiT360Plus
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.
ComfyUI nodes for DiT360 panoramic image generation with inpainting and outpainting support.
Looking for a different extension?
Custom Nodes (12)
README
ComfyUI-DiT360Plus
ComfyUI custom nodes for 360-degree panoramic image generation using DiT360 with FLUX.1-dev, plus an alternative Qwen-Image-2512 + Qwen 360 LoRA workflow. Supports text-to-panorama generation, inpainting, and outpainting.
Based on the research paper: DiT360: Panoramic Image Generation
Features
- Text-to-Panorama: Generate seamless 360 equirectangular panoramas from text prompts
- Inpainting: Edit specific regions of existing panoramic images
- Outpainting: Extend panoramic image boundaries with new content
- Smart VRAM Management: Four offload modes (off/model/balanced/sequential) with dynamic GPU allocation
- Edge Blending: Post-process for seamless horizontal wraparound
- 360 Preview: In-node panorama preview widget
Nodes
Pipeline
| Node | Description | |------|-------------| | Flux Panorama Loader | Loads any FLUX model (FLUX.1-dev, Kontext) with optional LoRA and configurable VRAM management | | DiT360 Text to Panorama | Generates 360 panorama from a text prompt | | DiT360 Pipeline Unloader | Frees GPU memory when done |
Editing
| Node | Description | |------|-------------| | Kontext Panorama Editor | Inpaints panorama regions using FLUX Kontext (image + mask + prompt) | | DiT360 Image Inverter | Inverts an image via RF-Inversion for editing | | DiT360 Panorama Editor | Inpaints or outpaints using inverted latents + mask |
Projection
| Node | Description | |------|-------------| | Equirect → Perspective | Extracts a distortion-free perspective view from an equirectangular panorama | | Perspective → Equirect | Composites an edited perspective patch back into the panorama with feathered blending |
Enhancement
| Node | Description | |------|-------------| | 360 Edge Blender | Blends left/right edges for seamless wrap | | 360 Empty Latent | Creates 2:1 aspect ratio latent | | DiT360 Mask Processor | Converts image/mask to binary editing mask | | 360 Viewer | Displays panorama preview in the node |
Installation
ComfyUI Manager (Recommended)
Search for ComfyUI-DiT360Plus in ComfyUI Manager and click Install.
Manual Installation
cd ComfyUI/custom_nodes
git clone https://github.com/thomashollier/ComfyUI-DiT360Plus.git
cd ComfyUI-DiT360Plus
pip install -r requirements.txt
Requirements
- Python 3.10+
- PyTorch 2.0+ (for
scaled_dot_product_attention) - CUDA GPU with 16GB+ VRAM (24GB recommended for 2048 width)
- FLUX.1-dev access on HuggingFace (requires
huggingface-cli login)
The DiT360 LoRA weights are downloaded automatically from Insta360-Research/DiT360-Panorama-Image-Generation.
Usage
Text to Panorama
The simplest workflow: generate a 360 panorama from text.
Flux Panorama Loader -> DiT360 Text to Panorama -> 360 Edge Blender -> 360 Viewer
- Add Flux Panorama Loader — select your FLUX model, set base_pipeline to FLUX.1-dev, add the DiT360 LoRA ID, choose an offload mode
- Add DiT360 Text to Panorama — enter your prompt (prefix with "This is a panorama image." for best results)
- Add 360 Edge Blender — ensures seamless horizontal wrapping
- Add 360 Viewer — preview the result
Recommended settings:
- Width: 2048 (height auto-calculated as 1024)
- Steps: 50
- Guidance scale: 2.8
- Seed: any
Kontext Inpainting
Edit specific regions of a panorama using FLUX Kontext — no inversion step needed.
Flux Panorama Loader (Kontext) -> Kontext Panorama Editor -> Save Image / 360 Viewer
Load Image (panorama + painted mask) --^
- Add Flux Panorama Loader — set base_pipeline to FLUX.1-Kontext-dev, leave LoRA empty
- Load your panorama and paint a mask on it (white = area to edit)
- Kontext Panorama Editor — enter your edit prompt, the masked region is repainted
Perspective Editing (Kontext)
Edit a local region of a panorama without equirectangular distortion artifacts. Extract a perspective view, edit it with Kontext, then composite back.
Load Image (panorama) → Equirect→Perspective → PreviewBridge → Kontext Panorama Editor → Perspective→Equirect → 360 Viewer
↗ (passthrough) ↗
Equirect→Perspective ─────────────────────────────────────────────────'
- Load your panorama and add Equirect → Perspective — set yaw/pitch to aim at the region, fov for zoom level
- Add PreviewBridge (from Impact Pack) — right-click → "Open in MaskEditor" to paint the edit mask on the perspective view
- Add Kontext Panorama Editor — enter your edit prompt, receives the perspective image + mask
- Add Perspective → Equirect — composites the edited patch back with feathered blending
- Add 360 Viewer to preview the result
Tip: Use ComfyUI-Impact-Pack for the PreviewBridge node, which lets you paint masks on images mid-workflow.
RF-Inversion Inpainting
Edit specific regions using the full RF-Inversion pipeline (more control, slower).
Flux Panorama Loader -> Image Inverter -----> Panorama Editor -> Edge Blender -> Viewer
Load Image (panorama) --^ ^
Load Image (mask) -> Mask Processor ------------'
- Load your panorama and a mask image (white = area to edit)
- DiT360 Image Inverter encodes the source image into editable latent space
- DiT360 Mask Processor converts the mask to the correct format
- DiT360 Panorama Editor (mode:
inpaint) regenerates the white mask region
Outpainting
Extend a panoramic image with new content.
Same workflow as RF-Inversion inpainting, but:
- The mask should have white = existing content to keep, black = area to generate
- Set mode to
outpaintin the Panorama Editor
VRAM Management
The Flux Panorama Loader offers four CPU offload strategies:
| Mode | Peak VRAM | Speed | Best For | |------|-----------|-------|----------| | off | ~24GB+ | Fastest | Large GPUs (48GB+) | | model | ~15GB | Fast | 1024 width on 24GB GPUs | | balanced | Configurable | Medium | 1536-2048 width on 24GB GPUs | | sequential | ~3GB | Slowest | Low VRAM GPUs (8-16GB) |
Balanced mode dynamically dispatches transformer layers to GPU after text encoding, maximizing GPU utilization based on available VRAM. The balanced_offload_gb parameter sets the maximum GPU budget — the system automatically reserves space for activations based on batch size and resolution.
Note: actual VRAM usage will be ~1GB above the set budget due to CUDA overhead, latent tensors, and prompt embeddings.
Key Parameters
| Parameter | Node | Description | Recommended |
|-----------|------|-------------|-------------|
| cpu_offload | FluxPanoramaLoader | VRAM management strategy | model for 1024, balanced for 2048 |
| balanced_offload_gb | FluxPanoramaLoader | GPU budget for balanced mode (GB) | 8-16 |
| guidance_scale | TextToPanorama, Editor | Classifier-free guidance | 2.8 |
| tau | Editor | Source preservation strength (0-100) | 50 |
| eta | Editor | Reconstruction vs. edit strength (lower = stronger edit) | 0.6-0.8 |
| mask_feather | Editor | Soft mask edge width as % of image width | 0-3 |
| gamma | Inverter | Inversion fidelity | 1.0 |
| blend_width | EdgeBlender | Edge blend width in pixels | 10-20 |
Alternative Model: Qwen-Image 360
In addition to FLUX.1-dev + DiT360, this project ships an example workflow for Qwen-Image-2512 + the Qwen 360 Diffusion LoRA. The Qwen path uses only built-in ComfyUI nodes (UNETLoader, CLIPLoader, VAELoader, KSampler, LoRA loaders) — no custom pipeline code — and the project's 360 Edge Blender + 360 Viewer are chained onto the output for seamless wrap and in-node preview.
Required Models
Place each file in the indicated ComfyUI/models/ subfolder:
| File | Folder | Source |
|------|--------|--------|
| qwen_image_2512_fp8_e4m3fn.safetensors | diffusion_models/ | Comfy-Org/Qwen-Image_ComfyUI |
| qwen_2.5_vl_7b_fp8_scaled.safetensors | text_encoders/ | Comfy-Org/Qwen-Image_ComfyUI |
| qwen_image_vae.safetensors | vae/ | Comfy-Org/Qwen-Image_ComfyUI |
| qwen-360-diffusion-2512-int8-bf16-v2.safetensors | loras/ | ProGamerGov/qwen-360-diffusion |
| Qwen-Image-2512-Lightning-4steps-V1.0-fp32.safetensors | loras/ | lightx2v/Qwen-Image-2512-Lightning |
Recommended Settings
- Resolution: 2048 × 1024 (2:1). Smaller 2:1 sizes work but horizons may be less stable.
- Sampler:
euler+simple,ModelSamplingAuraFlowshift = 3.1. - Lightning 4-step preset (default): 4 steps, CFG 1.0 — both LoRAs stacked.
- Full quality preset: 50 steps, CFG 4.0 — bypass the Lightning LoRA.
- Trigger phrases:
equirectangular,360 image,360 panorama, or360 degree panorama with equirectangular projection. The example workflow prefixes prompts withequirectangular 360 image, …. - People: describe head/face and shoes/boots for full-body shots to reduce limb distortion (poles of the sphere).
Notes
- The base model is FP8 while the 360 LoRA is int8-trained — rare patch/grid artifacts may appear. Lower the 360 LoRA strength or check the model card for int4 variants if you see them.
- The Lightning LoRA is a speed booster, not 360-specific; you can remove it for higher-quality slow runs.
Example Workflows
Example workflows are included in the examples/ folder:
text_to_panorama.json— Basic text-to-panorama generation (FLUX + DiT360)qwen_image_panorama.json— Text-to-panorama with Qwen-Image-2512 + Qwen 360 LoRA (Lightning 4-step)DiT360_inpainting.json— RF-Inversion inpainting/outpaintingLatLong_persp_inpainting.json— Perspective extraction + Kontext editing
Load these via ComfyUI's Load Workflow button.
How It Works
Circular Padding
DiT360 achieves seamless 360 wrapping by adding circular padding to the packed latent sequence. After FLUX packs spatial latents into a token sequence (B, N, D), the tokens are reshaped to a grid (B, n_h, n_w, D), and the first/last columns are wrapped:
[last_col | original_cols | first_col]
This lets the transformer "see" across the horizontal seam during denoising. The padding is removed after denoising before VAE decoding.
RF-Inversion (Editing)
Inpainting/outpainting uses RF-Inversion — a controlled forward/reverse ODE approach:
- Inversion: The source image is encoded through a forward ODE to get invertible noise representations (batch size 1, ~50 transformer passes)
- Editing: A controlled reverse ODE denoises with PersonalizeAnything attention processors that replace tokens at masked positions, preserving source content where specified (batch size 2, ~100 transformer passes)
Dynamic VRAM Management (Balanced Mode)
Balanced mode uses a phased approach to maximize GPU utilization:
- Text encoders load to GPU, encode prompt, offload to CPU (~10GB peak, brief)
- System queries actual free VRAM, estimates activation memory for the current batch size and resolution
accelerate.dispatch_modelpacks as many transformer layers onto GPU as the budget allows- Denoising runs with optimal GPU utilization
- Hooks removed, transformer back to CPU, VAE decodes
This avoids the VRAM spike that occurs when text encoders and transformer compete for GPU space simultaneously.
Credits
- DiT360 by Insta360 Research Team
- FLUX.1-dev by Black Forest Labs
- RF-Inversion by Rout et al.
- ComfyUI-DiT360 by cedarconnor (reference implementation)
License
Apache-2.0
Run ComfyUI workflows without the setup
No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.