Nodes/ComfyUI-DeZoomer-Nodes/Caption Refinement
ComfyUI Node

Caption Refinement

Refines and enhances captions using the Qwen2.5 model. Input Parameters: - caption: Input caption to refine (required) - system_prompt: Instructions for the model's behavior and output style - model_name: Qwen2.5 model to use (7B, 1.5B, or 72B variants) - temperature: Controls randomness in generation (0.1-1.0) - max_tokens: Maximum tokens for refinement output - quantization_type: Memory optimization (4-bit or 8-bit) - keep_model_loaded: Whether to keep the model in memory after processing - seed: Random seed for reproducible generation The node refines captions by: - Making the text more continuous and coherent - Removing video-specific references - Adding clothing details - Using only declarative sentences

By De-Zoomer·Created about a year ago·Updated about a year ago· 29
Caption Refinement
    • refined_caption
    caption
    system_promptYou are an AI prompt engineer tasked with helping me modifying a list of automatically generated prompts. Keep the original text but only do the following modifications: - you responses should just be the prompt - Write continuously, don't use multiple paragraphs, make the text form one coherent whole - do not mention your task or the text itself - remove references to video such as "the video begins" or "the video features" etc., but keep those sentences meaningful - mention the clothing details of the characters - use only declarative sentences
    model_nameQwen/Qwen2.5-7B-Instruct
    temperature0.7
    max_tokens200
    quantization_type4-bit
    keep_model_loadedfalse
    seed1
    CategoryDeZoomerNodes/text

    Inputs (8)

    NameTypeDefaultDescription
    captionSTRING
    system_promptSTRINGYou are an AI prompt engineer tasked with helping me modifying a list of automatically generated prompts. Keep the original text but only do the following modifications: - you responses should just be the prompt - Write continuously, don't use multiple paragraphs, make the text form one coherent whole - do not mention your task or the text itself - remove references to video such as "the video begins" or "the video features" etc., but keep those sentences meaningful - mention the clothing details of the characters - use only declarative sentences
    model_nameCOMBOQwen/Qwen2.5-7B-Instruct3 options: Qwen/Qwen2.5-7B-Instruct, Qwen/Qwen2.5-1.5B-Instruct, Qwen/Qwen2.5-72B-Instruct
    temperatureFLOAT0.70.1–1
    max_tokensINT20050–1000
    quantization_typeCOMBO4-bit2 options: 4-bit, 8-bit
    keep_model_loadedBOOLEANfalse
    seedINT11–18446744073709550000

    Outputs (1)

    NameTypeDescription
    refined_captionSTRING