ComfyUI Node
SparkDiffusion Text To Video
Text-to-video with the official SparkDiffusion inference (RoLA sparse attention + CrossDistill few-step sampling) in an isolated runtime.
SparkDiffusion Text To Video
- runtime
- profile
- video
- frames
- video_path
- metadata
◄promptA cat playing in the garden under the sun.►
◄seed0►
◄num_frames81►
◄aspect_ratio16:9►
◄filename_prefixsparkdiffusion/SparkWan►
◄load_framesfalse►
◄num_videos1►
CategorySparkDiffusion
Inputs (9)
| Name | Type | Default | Description |
|---|---|---|---|
| runtime | SPARKDIFFUSION_RUNTIME | — | |
| profile | SPARKDIFFUSION_PROFILE | — | |
| prompt | STRING | A cat playing in the garden under the sun. | — |
| seed | INT | 00–4294967295 | — |
| num_frames | INT | 815–161 | 4k+1 frames at 16 fps; 81 = 5 s (the length the checkpoints were distilled for). |
| aspect_ratio | COMBO | 16:9 | 5 options: 16:9, 9:16, 1:1, 4:3, 3:4 |
| filename_prefix | STRING | sparkdiffusion/SparkWan | — |
| load_framesopt | BOOLEAN | false | Decode every frame into the IMAGE output (memory heavy). Off = first frame only. |
| num_videosopt | INT | 11–16 | Videos generated in ONE runtime process with seeds seed, seed+1, ... (upstream NUM_SAMPLES). Loading and torch.compile are paid once, so extra videos are cheap. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| video | VIDEO | — |
| frames | IMAGE | — |
| video_path | STRING | — |
| metadata | STRING | — |