Nodes/ComfyUI-SparkDiffusion/SparkDiffusion Text To Video
ComfyUI Node

SparkDiffusion Text To Video

Text-to-video with the official SparkDiffusion inference (RoLA sparse attention + CrossDistill few-step sampling) in an isolated runtime.

By hiroki-abe-58·Created 4 days ago·Updated 4 days ago· 1
SparkDiffusion Text To Video
  • runtime
  • profile
  • video
  • frames
  • video_path
  • metadata
◄promptA cat playing in the garden under the sun.►
◄seed0►
◄num_frames81►
◄aspect_ratio16:9►
◄filename_prefixsparkdiffusion/SparkWan►
◄load_framesfalse►
◄num_videos1►
CategorySparkDiffusion

Inputs (9)

NameTypeDefaultDescription
runtimeSPARKDIFFUSION_RUNTIME—
profileSPARKDIFFUSION_PROFILE—
promptSTRINGA cat playing in the garden under the sun.—
seedINT00–4294967295—
num_framesINT815–1614k+1 frames at 16 fps; 81 = 5 s (the length the checkpoints were distilled for).
aspect_ratioCOMBO16:95 options: 16:9, 9:16, 1:1, 4:3, 3:4
filename_prefixSTRINGsparkdiffusion/SparkWan—
load_framesoptBOOLEANfalseDecode every frame into the IMAGE output (memory heavy). Off = first frame only.
num_videosoptINT11–16Videos generated in ONE runtime process with seeds seed, seed+1, ... (upstream NUM_SAMPLES). Loading and torch.compile are paid once, so extra videos are cheap.

Outputs (4)

NameTypeDescription
videoVIDEO—
framesIMAGE—
video_pathSTRING—
metadataSTRING—