Nodes/ComfyUI-BytePlus-ModelArk/BytePlus Seedance Image to Video
ComfyUI Node

BytePlus Seedance Image to Video

One still, six seconds of life

By byteplus-sa·Created 8 days ago·Updated about 8 hours ago· 3
BytePlus Seedance Image to Video
  • image
  • VIDEO
  • last_frame
  • response
◄modelseedance-1-0-pro-fast-251015►
◄prompt—►
◄resolution▾►
◄aspect_ratio▾►
◄duration5►
◄seed0►
◄camera_fixedfalse►
◄watermarkfalse►
◄enable_offline_inferencefalse►
◄generation_count1►
◄non_blockingfalse►

Text-to-video makes you describe a whole shot and hope. Image-to-video starts from a picture you've already approved, which is why it's the node most people reach for first: you get your composition from a model you trust (Seedream, a local checkpoint, a photo), and you use Seedance only for motion. The pack even ships a Text to Image to Video template doing exactly that - Seedream paints the first frame, this node animates it.

That division of labour is the real argument. Pixel quality comes from the image model, motion comes from the video model, and neither has to be good at the other's job.

How it works

image and prompt are both required. The image is the first frame, sent inline as base64 - no upload, no Comfy.org login, nothing to manage. The prompt describes motion: what moves, how the camera behaves, what changes over the six seconds. Re-describing what's already visible in the frame mostly wastes tokens.

  • model is a plain picker: seedance-1-0-pro-fast (the default) or seedance-1-0-pro. Fast is the one you'll use for iteration.
  • resolution runs 480p, 720p, 1080p. aspect_ratio includes adaptive alongside the standard ratios, and adaptive is what you want - it follows the input image's framing instead of cropping your carefully composed shot into 16:9.
  • duration is 2 to 12 seconds, default 5.
  • camera_fixed appends a fixed-camera instruction to your prompt. It's a nudge, not a guarantee - BytePlus says so in the tooltip - but on a portrait or product shot it noticeably reduces the slow zoom-drift these models love.
  • enable_offline_inference is the flex tier: cheaper, but results arrive within 48 hours. Great for overnight batch work, terrible for iterating.
  • generation_count and non_blocking are the pack's extras. The first fires parallel generations (each on seed + N, each billed); the second submits and returns so you can collect the finished video on a later run.
  • seed is a re-run trigger, not a reproducibility control.

Outputs: VIDEO - a list when you generate more than one, in which case the downstream node runs once per video - plus last_frame, the final frames as one image batch, and response, the task JSON. Feed last_frame into another Image to Video node and you've built a crude shot extender; the pack's own extension template does the same thing more carefully on the 2.5 side.

Install and key

cd ComfyUI/custom_nodes
git clone https://github.com/byteplus-sa/ComfyUI-BytePlus-ModelArk
pip install -r ComfyUI-BytePlus-ModelArk/requirements.txt

Restart (ComfyUI 0.31.0 or newer), or install from Manager by searching BytePlus ModelArk - and grab the example workflows while you're there, since the Text to Image to Video template is a better starting point than a blank canvas. Save the ModelArk API key and region in Settings → BytePlus, or BYTEPLUS_API_KEY / BYTEPLUS_REGION in user/.env, and enable the model in the console.

The node saves nothing. Connect a Save Video; an unwired node doesn't run and isn't billed.

Where people get burned

Keep the aspect ratio in mind. If your input image is 9:16 and you set aspect_ratio to 16:9, you're asking the model to re-frame the shot while animating it - sometimes fine, often a slow reveal of cropped nonsense. adaptive exists so you don't have to think about it.

The other recurring one is motion quality against duration. There's a reason BytePlus caps 1.0 at 12 seconds: the longer the clip, the more the model has to invent, and the more likely you get warping where a second ago there was a face. Six seconds of a small, deliberate movement looks better than twelve seconds of drifting. If you need length, generate two clips and join them, or step up to the 2.5 nodes where 30 seconds is supported.

And the standard cost reminder: video is metered per second of output. There's no cheaper way to learn your prompt was wrong than image-to-video on a low resolution, which is why the workflow in this pack is image first, motion second.

CategoryBytePlus ModelArk

Inputs (12)

NameTypeDefaultDescription
modelCOMBOseedance-1-0-pro-fast-2510152 options: seedance-1-0-pro-250528, seedance-1-0-pro-fast-251015
promptSTRINGThe text prompt used to generate the video.
imageIMAGEFirst frame to be used for the video.
resolutionCOMBOThe resolution of the output video.
aspect_ratioCOMBOThe aspect ratio of the output video.
durationINT52–12The duration of the output video in seconds.
seedoptINT00–2147483647Seed to use for generation.
camera_fixedoptBOOLEANfalseSpecifies whether to fix the camera. The platform appends an instruction to fix the camera to your prompt, but does not guarantee the actual effect.
watermarkoptBOOLEANfalseWhether to add an "AI generated" watermark to the video.
enable_offline_inferenceoptBOOLEANfalseUse the flex (offline) service tier: lower price, results within 48 hours.
generation_countoptINT1Number of separate generations to run in parallel. With several, generation N uses seed + N so the results differ.
non_blockingoptBOOLEANfalseSubmit the task and return at once; run the node again to collect the finished video.

Outputs (3)

NameTypeDescription
VIDEOVIDEOThe generated video, or every video of a generation_count batch (the next node runs once per video).
last_frameIMAGELast frame of each generated video, as one image batch in the same order as the videos.
responseSTRINGTask results as JSON, or the pending task IDs of a non_blocking run.