ComfyUI Node: Qwen Omni Combined🐼

Authored by SXQBW

Created

Updated

41 stars

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Category

🐼QwenOmni

Inputs

model_name
  • Qwen2.5-Omni-3B
  • Qwen2.5-Omni-7B
quantization
  • 👍 4-bit (VRAM-friendly)
  • ⚖️ 8-bit (Balanced Precision)
  • 🚫 None (Original Precision)
prompt STRING
audio_output
  • 🔇None (No Audio)
  • 👱‍♀️Chelsie (Female)
  • 👨‍🦰Ethan (Male)
audio_source
  • 🎧 Separate Audio Input
  • 🎬 Video Built-in Audio
max_tokens INT
temperature FLOAT
top_p FLOAT
repetition_penalty FLOAT
image IMAGE
audio AUDIO
video_path VIDEO_PATH

Outputs

STRING

AUDIO

Extension: ComfyUI-Qwen-Omni

ComfyUI-Qwen-Omni is the first ComfyUI plugin that supports end-to-end multimodal interaction, enabling seamless joint generation and editing of text, images, and audio. Without intermediate steps, with just one operation, the model can simultaneously understand and process multiple input modalities, generating coherent text descriptions and voice outputs, providing an unprecedentedly smooth experience for AI creation.

Authored by SXQBW

Looking for a different node?

More nodes in ComfyUI-Qwen-Omni

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Learn more