Nodes/ComfyUI/Vidu Q3 Text-to-Video Generation
ComfyUI Node Runs on cloud

Vidu Q3 Text-to-Video Generation

The one people actually talk about

By Comfy-Org·Created 4 years ago·Updated about 15 hours ago· 130,663
Vidu Q3 Text-to-Video Generation
    • VIDEO
    model
    prompt
    seed1

    If you've heard anything about Vidu from the community, it was probably about Q3. The r/comfyui thread that got Vidu its Western reputation was a solo animator shipping an AI sci-fi series and posting clips where "Vidu Q3 is nailing my character expressions" - subtle eye and ear movements that local models like Wan 2.2 and LTX 2.3 still miss. This node is that model as a text-to-video node, and it's the one Vidu node with a genuine quality story to tell rather than just a price tag.

    Two models on offer, and they're the same split as the rest of the Q3 family:

    • viduq3-pro - the flagship. Better expressions, better motion, higher cost.
    • viduq3-turbo - the fast lane. Same scene-building, less fidelity on the details that make Q3 famous. Iterate here, ship on pro.

    What Q3 gives you that older Vidu didn't

    The headline feature is audio. Both models carry an audio toggle - turn it on and the output includes dialogue and sound effects, not just a silent clip. That's a real capability jump: most of the Vidu 1 and 2 nodes are silent by design. It's also a cost multiplier, so the toggle is where the price badge grows.

    The other upgrade is duration - 1 to 16 seconds, the longest native range in the Vidu family. Combined with audio, that makes Q3 the "actually produce a scene" node: 12 seconds of a character talking with voice is a chunk of usable footage, not a flash of motion.

    The rest of the inputs are familiar: prompt (required, max 2000 chars), aspect_ratio (16:9, 9:16, 3:4, 4:3, 1:1), resolution (720p/1080p), and seed (defaults to 1, 0 for random). Note that model-specific settings like duration, resolution, and the audio toggle live inside the model dropdown - pick the model and its options expand beneath it. That's the dynamic-combo pattern ComfyUI uses across these partner nodes, and it's the thing that trips people up the first time: the settings you're looking for aren't missing, they're nested in the model widget.

    How it runs

    ComfyUI POSTs to Vidu's text-to-video endpoint through the Comfy proxy, polls the task, and downloads the clip. Needs a Comfy account with credits, internet, no local GPU. Output is a VIDEO object - save it downstream.

    The honest version

    Q3 is the model that justifies Vidu's existence, but it's also the expensive one, and the community says so plainly - the thread praising Q3's expressions has people asking "how much is this costing you per episode?" and the answer is "a lot." The cost-aware workflow is: compose and iterate at 720p on turbo with audio off, lock the seed, then spend on a pro 1080p run with audio. And remember the standing rule: ComfyUI gives you the pipeline, not looser moderation - Q3's content rules are the same here as on Vidu's own site. If you want the expressions without the bill, you're asking for an open model, and nobody has one that does it yet.

    Categorypartner/video/Vidu

    Inputs (3)

    NameTypeDefaultDescription
    modelCOMBOModel to use for video generation.
    promptSTRINGA textual description for video generation, with a maximum length of 2000 characters.
    seedINT10–2147483647

    Outputs (1)

    NameTypeDescription
    VIDEOVIDEO