Nodes/ComfyUI/PixVerse Text to Video
ComfyUI Node Runs on cloud

PixVerse Text to Video

Prompt-only clips with a motion-mode switch

By Comfy-Org·Created 4 years ago·Updated about an hour ago· 129,820
PixVerse Text to Video
  • pixverse_template
  • VIDEO
prompt
aspect_ratio
quality540p
duration_seconds
motion_mode
seed0
negative_prompt

If you just want to type a sentence and get a video without fiddling with a start frame, PixVerse's text-to-video node is the no-input-media version of its image-to-video sibling - and it's built into ComfyUI core as a partner/video API node. No install, nothing runs on your GPU; the generation happens on PixVerse's servers through Comfy's proxy, billed per render from your Comfy account credits.

The input list is short: prompt (the whole creative load lives here), aspect_ratio (16:9 down to 9:16 with the usual stops in between), quality defaulting to 540p with a ceiling of 1080p, duration_seconds (5 or 8), motion_mode (normal or fast), and seed as the standard re-run switch. Optional extras: negative_prompt, and pixverse_template - a PIXVERSE_TEMPLATE input from the sibling PixVerse Template node that injects a style preset.

The constraints are where you need the honest version, and they mirror the image-to-video node exactly. Choose 1080p and PixVerse silently pins you to 5 seconds and normal motion - the higher resolution simply doesn't support the fast tier or the longer duration. Choose the 8-second duration and motion is forced to normal. Fast motion means 5 seconds at ≤720p, full stop. If your clip comes back feeling sedate, the first thing to check isn't your prompt - it's whether you're even allowed to be in fast mode.

motion_mode is the one creative decision with real teeth. normal is the grounded, natural-motion default; fast is PixVerse's higher-energy mode, better for stylized and exaggerated animation. Since the model has no reference image to anchor on, the prompt has to carry the whole world - subjects, setting, camera, and motion style - and it's worth writing like you're directing a shot rather than describing a scene. The node is lenient about empty-ish prompts (it only requires a non-empty string), but "cinematic" alone will get you a video that looks exactly as cheap as that sounds.

The generation is async and can sit in the queue for tens of seconds up to a couple of minutes while the node polls, so budget for that in a long queue of jobs. Output is a single VIDEO that drops into Save/Preview or downstream nodes like every other node in the family.

Honest positioning: this is the quick, low-commitment end of the PixVerse family. No start frame means less control over composition and less subject consistency, and the 5–8 second / ≤1080p envelope is a real ceiling. Use it for drafts, ideation, and any shot where you don't have an image to start from - and when you need the output to actually match something, move up to the image-to-video node and feed it a frame.

The usual partner-node reminders apply: your prompt is sent to PixVerse's servers, their content moderation can reject a generation (that failure is the filter, not your graph), and the price badge is there to remind you that drafting at 540p costs far less than iterating at 1080p. PixVerse's API nodes have been in core since mid-2025, and this one remains the simplest way to get a hosted PixVerse clip from a sentence.

Categorypartner/video/PixVerse

Inputs (8)

NameTypeDefaultDescription
promptSTRINGPrompt for the video generation
aspect_ratioCOMBO5 options: 16:9, 4:3, 1:1, 3:4, 9:16
qualityCOMBO540p4 options: 360p, 540p, 720p, 1080p
duration_secondsCOMBO2 options: 5, 8
motion_modeCOMBO2 options: normal, fast
seedINT00–2147483647Seed for video generation.
negative_promptoptSTRINGAn optional text description of undesired elements on an image.
pixverse_templateoptPIXVERSE_TEMPLATEAn optional template to influence style of generation, created by the PixVerse Template node.

Outputs (1)

NameTypeDescription
VIDEOVIDEO