ComfyUI Node: ThinkingLLM Gemma 4 Audio (GGUF)

Authored by goodguy1963

Created

Updated

13 stars

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Category

ThinkingLLM

Inputs

model_name
  • gemma-4-12b-it-BF16.gguf [~24GB]
  • gemma-4-12b-it-Q4_K_M.gguf [~7.5GB]
  • gemma-4-12b-it-Q4_K_S.gguf [~7.5GB]
  • gemma-4-12b-it-Q5_K_M.gguf [~9GB]
  • gemma-4-12b-it-Q5_K_S.gguf [~9GB]
  • gemma-4-12b-it-Q6_K.gguf [~11GB]
  • gemma-4-12b-it-Q8_0.gguf [~13GB]
  • gemma-4-12b-it-UD-Q4_K_XL.gguf
  • google_gemma-4-E4B-it-Q4_K_M.gguf [~5.4GB]
  • google_gemma-4-E4B-it-Q5_K_M.gguf [~5.8GB]
  • google_gemma-4-E4B-it-Q6_K.gguf [~6.3GB]
  • google_gemma-4-E4B-it-Q8_0.gguf [~8GB]
  • google_gemma-4-E4B-it-bf16.gguf [~15GB]
custom_prompt STRING
max_tokens INT
keep_model_loaded BOOLEAN
seed INT
stream_tokens_to_terminal BOOLEAN
enable_thinking BOOLEAN
auto_finalization_retry BOOLEAN
hf_token STRING
audio AUDIO
audio_file_path STRING

Outputs

STRING

STRING

Extension: ComfyUI-ThinkingLLM

A multimodal ComfyUI AI node with Qwen3.5, Qwen3-VL, Qwen2.5-VL, Qwen3, and Gemma 4 integrations. Features live thinking in the terminal to see what the LLM is doing in real time.

Authored by goodguy1963

Looking for a different node?

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Learn more