BytePlus Seedance Image to Video
One still, six seconds of life
- image
- VIDEO
- last_frame
- response
Text-to-video makes you describe a whole shot and hope. Image-to-video starts from a picture you've already approved, which is why it's the node most people reach for first: you get your composition from a model you trust (Seedream, a local checkpoint, a photo), and you use Seedance only for motion. The pack even ships a Text to Image to Video template doing exactly that - Seedream paints the first frame, this node animates it.
That division of labour is the real argument. Pixel quality comes from the image model, motion comes from the video model, and neither has to be good at the other's job.
How it works
image and prompt are both required. The image is the first frame, sent inline as base64 - no upload, no Comfy.org login, nothing to manage. The prompt describes motion: what moves, how the camera behaves, what changes over the six seconds. Re-describing what's already visible in the frame mostly wastes tokens.
modelis a plain picker:seedance-1-0-pro-fast(the default) orseedance-1-0-pro. Fast is the one you'll use for iteration.resolutionruns 480p, 720p, 1080p.aspect_ratioincludesadaptivealongside the standard ratios, andadaptiveis what you want - it follows the input image's framing instead of cropping your carefully composed shot into 16:9.durationis 2 to 12 seconds, default 5.camera_fixedappends a fixed-camera instruction to your prompt. It's a nudge, not a guarantee - BytePlus says so in the tooltip - but on a portrait or product shot it noticeably reduces the slow zoom-drift these models love.enable_offline_inferenceis the flex tier: cheaper, but results arrive within 48 hours. Great for overnight batch work, terrible for iterating.generation_countandnon_blockingare the pack's extras. The first fires parallel generations (each onseed + N, each billed); the second submits and returns so you can collect the finished video on a later run.seedis a re-run trigger, not a reproducibility control.
Outputs: VIDEO - a list when you generate more than one, in which case the downstream node runs once per video - plus last_frame, the final frames as one image batch, and response, the task JSON. Feed last_frame into another Image to Video node and you've built a crude shot extender; the pack's own extension template does the same thing more carefully on the 2.5 side.
Install and key
cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt
Restart (ComfyUI 0.31.0 or newer), or install from Manager by searching BytePlus ModelArk - and grab the example workflows while you're there, since the Text to Image to Video template is a better starting point than a blank canvas. Save the ModelArk API key and region in Settings → BytePlus, or BYTEPLUS_API_KEY / BYTEPLUS_REGION in user/.env, and enable the model in the console.
The node saves nothing. Connect a Save Video; an unwired node doesn't run and isn't billed.
Where people get burned
Keep the aspect ratio in mind. If your input image is 9:16 and you set aspect_ratio to 16:9, you're asking the model to re-frame the shot while animating it - sometimes fine, often a slow reveal of cropped nonsense. adaptive exists so you don't have to think about it.
The other recurring one is motion quality against duration. There's a reason BytePlus caps 1.0 at 12 seconds: the longer the clip, the more the model has to invent, and the more likely you get warping where a second ago there was a face. Six seconds of a small, deliberate movement looks better than twelve seconds of drifting. If you need length, generate two clips and join them, or step up to the 2.5 nodes where 30 seconds is supported.
And the standard cost reminder: video is metered per second of output. There's no cheaper way to learn your prompt was wrong than image-to-video on a low resolution, which is why the workflow in this pack is image first, motion second.
Inputs (12)
| Name | Type | Default | Description |
|---|---|---|---|
| model | COMBO | seedance-1-0-pro-fast-251015 | 2 options: seedance-1-0-pro-250528, seedance-1-0-pro-fast-251015 |
| prompt | STRING | The text prompt used to generate the video. | |
| image | IMAGE | First frame to be used for the video. | |
| resolution | COMBO | The resolution of the output video. | |
| aspect_ratio | COMBO | The aspect ratio of the output video. | |
| duration | INT | 52–12 | The duration of the output video in seconds. |
| seedopt | INT | 00–2147483647 | Seed to use for generation. |
| camera_fixedopt | BOOLEAN | false | Specifies whether to fix the camera. The platform appends an instruction to fix the camera to your prompt, but does not guarantee the actual effect. |
| watermarkopt | BOOLEAN | false | Whether to add an "AI generated" watermark to the video. |
| enable_offline_inferenceopt | BOOLEAN | false | Use the flex (offline) service tier: lower price, results within 48 hours. |
| generation_countopt | INT | 1 | Number of separate generations to run in parallel. With several, generation N uses seed + N so the results differ. |
| non_blockingopt | BOOLEAN | false | Submit the task and return at once; run the node again to collect the finished video. |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| VIDEO | VIDEO | The generated video, or every video of a generation_count batch (the next node runs once per video). |
| last_frame | IMAGE | Last frame of each generated video, as one image batch in the same order as the videos. |
| response | STRING | Task results as JSON, or the pending task IDs of a non_blocking run. |