Nodes/ComfyUI-Qwen-Omni/Qwen Omni Combined🐼
ComfyUI Node

Qwen Omni Combined🐼

A ComfyUI node in 🐼QwenOmni with 12 inputs and 2 outputs.

By SXQBW·Created about a year ago·Updated about a year ago· 41
Qwen Omni Combined🐼
  • image
  • audio
  • video_path
  • text
  • audio
model_nameQwen2.5-Omni-3B
quantization👍 4-bit (VRAM-friendly)
promptHi!😽
audio_output🔇None (No Audio)
audio_source🎧 Separate Audio Input
max_tokens132
temperature0.4
top_p0.90
repetition_penalty1.00
Category🐼QwenOmni

Inputs (12)

NameTypeDefaultDescription
model_nameCOMBOQwen2.5-Omni-3BSelect the available model version. 选择可用的模型版本。
quantizationCOMBO👍 4-bit (VRAM-friendly)Select the quantization level: ✅ 4-bit: Significantly reduces VRAM usage, suitable for resource-constrained environments. ⚖️ 8-bit: Strikes a balance between precision and performance. 🚫 None: Uses the original floating-point precision (requires a high-end GPU). 选择量化级别: ✅ 4位:显著减少VRAM使用,适合资源受限环境。 ⚖️ 8位:在精度和性能之间取得平衡。 🚫 无:使用原始浮点精度(需要高端GPU)。
promptSTRINGHi!😽Enter a text prompt, supporting Chinese and emojis. Example: 'Describe a cat in a painter's style.' 输入文本提示,支持中文和表情符号。示例:'以画家的风格描述一只猫。'
audio_outputCOMBO🔇None (No Audio)Audio output options: 🔇 Do not generate audio. 👱‍♀️ Use the female voice Chelsie (warm tone). 👨‍🦰 Use the male voice Ethan (calm tone). 音频输出选项: 🔇 不生成音频。 👱‍♀️ 使用女性声音Chelsie(温暖语调)。 👨‍🦰 使用男性声音Ethan(平静语调)。
audio_sourceCOMBO🎧 Separate Audio InputSelect audio source: Use video's built-in audio track (priority) / Input a separate audio file (external audio) 选择音频源:使用视频内置音轨(优先)/输入单独的音频文件(外部音频)
max_tokensINT13264–2048Control the maximum length of the generated text (in tokens). Generally, 100 tokens correspond to approximately 50 - 100 Chinese characters or 67 - 100 English words, but the actual number may vary depending on the text content and the model's tokenization strategy. Recommended range: 64 - 512. 控制生成文本的最大长度(以token为单位)。 一般来说,100个token约对应50-100个汉字或67-100个英文单词,但实际数量可能因文本内容和模型的分词策略而异。 推荐范围:64-512。
temperatureFLOAT0.40.1–1Control the generation diversity: ▫️ 0.1 - 0.3: Generate structured/technical content. ▫️ 0.5 - 0.7: Balance creativity and logic. ▫️ 0.8 - 1.0: High degree of freedom (may produce incoherent content). 控制生成多样性: ▫️ 0.1-0.3:生成结构化/技术性内容。 ▫️ 0.5-0.7:平衡创造性和逻辑性。 ▫️ 0.8-1.0:高度自由(可能产生不连贯内容)。
top_pFLOAT0.900–1Nucleus sampling threshold: ▪️ Close to 1.0: Retain more candidate words (more random). ▪️ 0.5 - 0.8: Balance quality and diversity. ▪️ Below 0.3: Generate more conservative content. 核采样阈值: ▪️ 接近1.0:保留更多候选词(更随机)。 ▪️ 0.5-0.8:平衡质量和多样性。 ▪️ 低于0.3:生成更保守的内容。
repetition_penaltyFLOAT1.000–2Control of repeated content: ⚠️ 1.0: Default behavior. ⚠️ >1.0 (Recommended 1.2): Suppress repeated phrases. ⚠️ <1.0 (Recommended 0.8): Encourage repeated emphasis. 控制重复内容: ⚠️ 1.0:默认行为。 ⚠️ >1.0(推荐1.2):抑制重复短语。 ⚠️ <1.0(推荐0.8):鼓励重复强调。
imageoptIMAGEUpload a reference image (supports PNG/JPG), and the model will adjust the generation result based on the image content. 上传参考图像(支持PNG/JPG),模型将根据图像内容调整生成结果。
audiooptAUDIOUpload an audio file (supports MP3/WAV), and the model will analyze the audio content and generate relevant responses. 上传音频文件(支持MP3/WAV),模型将分析音频内容并生成相关响应。
video_pathoptVIDEO_PATHEnter the video file (supports MP4/WEBM), and the model will extract visual features to assist in generation. 输入视频文件路径(支持MP4/WEBM),模型将提取视觉特征辅助生成。

Outputs (2)

NameTypeDescription
textSTRING
audioAUDIO