ComfyUI Node: JoyCaption

Authored by fpgaminer

Created

Updated

121 stars

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Category

JoyCaption

Inputs

model JOYCAPMODEL
image IMAGE
caption_type
  • Descriptive
  • Descriptive (Casual)
  • Straightforward
  • Stable Diffusion Prompt
  • MidJourney
  • Danbooru tag list
  • e621 tag list
  • Rule34 tag list
  • Booru-like tag list
  • Art Critic
  • Product Listing
  • Social Media Post
caption_length
  • any
  • very short
  • short
  • medium-length
  • long
  • very long
  • 20
  • 30
  • 40
  • 50
  • 60
  • 70
  • 80
  • 90
  • 100
  • 110
  • 120
  • 130
  • 140
  • 150
  • 160
  • 170
  • 180
  • 190
  • 200
  • 210
  • 220
  • 230
  • 240
  • 250
  • 260
extra_option1
  • If there is a person/character in the image you must refer to them as {name}.
  • Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
  • Include information about lighting.
  • Include information about camera angle.
  • Include information about whether there is a watermark or not.
  • Include information about whether there are JPEG artifacts or not.
  • If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
  • Do NOT include anything sexual; keep it PG.
  • Do NOT mention the image's resolution.
  • You MUST include information about the subjective aesthetic quality of the image from low to very high.
  • Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
  • Do NOT mention any text that is in the image.
  • Specify the depth of field and whether the background is in focus or blurred.
  • If applicable, mention the likely use of artificial or natural lighting sources.
  • Do NOT use any ambiguous language.
  • Include whether the image is sfw, suggestive, or nsfw.
  • ONLY describe the most important elements of the image.
  • If it is a work of art, do not include the artist's name or the title of the work.
  • Identify the image orientation (portrait, landscape, or square) and aspect ratio if obvious.
  • Use vulgar slang and profanity, such as (but not limited to) "fucking," "slut," "cock," etc.
  • Do NOT use polite euphemisms—lean into blunt, casual phrasing.
  • Include information about the ages of any people/characters when applicable.
  • Mention whether the image depicts an extreme close-up, close-up, medium close-up, medium shot, cowboy shot, medium wide shot, wide shot, or extreme wide shot.
  • Do not mention the mood/feeling/etc of the image.
  • Explicitly specify the vantage height (eye-level, low-angle worm’s-eye, bird’s-eye, drone, rooftop, etc.).
  • If there is a watermark, you must mention it.
  • Your response will be used by a text-to-image model, so avoid useless meta phrases like “This image shows…”, "You are looking at...", etc.
extra_option2
  • If there is a person/character in the image you must refer to them as {name}.
  • Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
  • Include information about lighting.
  • Include information about camera angle.
  • Include information about whether there is a watermark or not.
  • Include information about whether there are JPEG artifacts or not.
  • If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
  • Do NOT include anything sexual; keep it PG.
  • Do NOT mention the image's resolution.
  • You MUST include information about the subjective aesthetic quality of the image from low to very high.
  • Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
  • Do NOT mention any text that is in the image.
  • Specify the depth of field and whether the background is in focus or blurred.
  • If applicable, mention the likely use of artificial or natural lighting sources.
  • Do NOT use any ambiguous language.
  • Include whether the image is sfw, suggestive, or nsfw.
  • ONLY describe the most important elements of the image.
  • If it is a work of art, do not include the artist's name or the title of the work.
  • Identify the image orientation (portrait, landscape, or square) and aspect ratio if obvious.
  • Use vulgar slang and profanity, such as (but not limited to) "fucking," "slut," "cock," etc.
  • Do NOT use polite euphemisms—lean into blunt, casual phrasing.
  • Include information about the ages of any people/characters when applicable.
  • Mention whether the image depicts an extreme close-up, close-up, medium close-up, medium shot, cowboy shot, medium wide shot, wide shot, or extreme wide shot.
  • Do not mention the mood/feeling/etc of the image.
  • Explicitly specify the vantage height (eye-level, low-angle worm’s-eye, bird’s-eye, drone, rooftop, etc.).
  • If there is a watermark, you must mention it.
  • Your response will be used by a text-to-image model, so avoid useless meta phrases like “This image shows…”, "You are looking at...", etc.
extra_option3
  • If there is a person/character in the image you must refer to them as {name}.
  • Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
  • Include information about lighting.
  • Include information about camera angle.
  • Include information about whether there is a watermark or not.
  • Include information about whether there are JPEG artifacts or not.
  • If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
  • Do NOT include anything sexual; keep it PG.
  • Do NOT mention the image's resolution.
  • You MUST include information about the subjective aesthetic quality of the image from low to very high.
  • Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
  • Do NOT mention any text that is in the image.
  • Specify the depth of field and whether the background is in focus or blurred.
  • If applicable, mention the likely use of artificial or natural lighting sources.
  • Do NOT use any ambiguous language.
  • Include whether the image is sfw, suggestive, or nsfw.
  • ONLY describe the most important elements of the image.
  • If it is a work of art, do not include the artist's name or the title of the work.
  • Identify the image orientation (portrait, landscape, or square) and aspect ratio if obvious.
  • Use vulgar slang and profanity, such as (but not limited to) "fucking," "slut," "cock," etc.
  • Do NOT use polite euphemisms—lean into blunt, casual phrasing.
  • Include information about the ages of any people/characters when applicable.
  • Mention whether the image depicts an extreme close-up, close-up, medium close-up, medium shot, cowboy shot, medium wide shot, wide shot, or extreme wide shot.
  • Do not mention the mood/feeling/etc of the image.
  • Explicitly specify the vantage height (eye-level, low-angle worm’s-eye, bird’s-eye, drone, rooftop, etc.).
  • If there is a watermark, you must mention it.
  • Your response will be used by a text-to-image model, so avoid useless meta phrases like “This image shows…”, "You are looking at...", etc.
extra_option4
  • If there is a person/character in the image you must refer to them as {name}.
  • Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
  • Include information about lighting.
  • Include information about camera angle.
  • Include information about whether there is a watermark or not.
  • Include information about whether there are JPEG artifacts or not.
  • If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
  • Do NOT include anything sexual; keep it PG.
  • Do NOT mention the image's resolution.
  • You MUST include information about the subjective aesthetic quality of the image from low to very high.
  • Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
  • Do NOT mention any text that is in the image.
  • Specify the depth of field and whether the background is in focus or blurred.
  • If applicable, mention the likely use of artificial or natural lighting sources.
  • Do NOT use any ambiguous language.
  • Include whether the image is sfw, suggestive, or nsfw.
  • ONLY describe the most important elements of the image.
  • If it is a work of art, do not include the artist's name or the title of the work.
  • Identify the image orientation (portrait, landscape, or square) and aspect ratio if obvious.
  • Use vulgar slang and profanity, such as (but not limited to) "fucking," "slut," "cock," etc.
  • Do NOT use polite euphemisms—lean into blunt, casual phrasing.
  • Include information about the ages of any people/characters when applicable.
  • Mention whether the image depicts an extreme close-up, close-up, medium close-up, medium shot, cowboy shot, medium wide shot, wide shot, or extreme wide shot.
  • Do not mention the mood/feeling/etc of the image.
  • Explicitly specify the vantage height (eye-level, low-angle worm’s-eye, bird’s-eye, drone, rooftop, etc.).
  • If there is a watermark, you must mention it.
  • Your response will be used by a text-to-image model, so avoid useless meta phrases like “This image shows…”, "You are looking at...", etc.
extra_option5
  • If there is a person/character in the image you must refer to them as {name}.
  • Do NOT include information about people/characters that cannot be changed (like ethnicity, gender, etc), but do still include changeable attributes (like hair style).
  • Include information about lighting.
  • Include information about camera angle.
  • Include information about whether there is a watermark or not.
  • Include information about whether there are JPEG artifacts or not.
  • If it is a photo you MUST include information about what camera was likely used and details such as aperture, shutter speed, ISO, etc.
  • Do NOT include anything sexual; keep it PG.
  • Do NOT mention the image's resolution.
  • You MUST include information about the subjective aesthetic quality of the image from low to very high.
  • Include information on the image's composition style, such as leading lines, rule of thirds, or symmetry.
  • Do NOT mention any text that is in the image.
  • Specify the depth of field and whether the background is in focus or blurred.
  • If applicable, mention the likely use of artificial or natural lighting sources.
  • Do NOT use any ambiguous language.
  • Include whether the image is sfw, suggestive, or nsfw.
  • ONLY describe the most important elements of the image.
  • If it is a work of art, do not include the artist's name or the title of the work.
  • Identify the image orientation (portrait, landscape, or square) and aspect ratio if obvious.
  • Use vulgar slang and profanity, such as (but not limited to) "fucking," "slut," "cock," etc.
  • Do NOT use polite euphemisms—lean into blunt, casual phrasing.
  • Include information about the ages of any people/characters when applicable.
  • Mention whether the image depicts an extreme close-up, close-up, medium close-up, medium shot, cowboy shot, medium wide shot, wide shot, or extreme wide shot.
  • Do not mention the mood/feeling/etc of the image.
  • Explicitly specify the vantage height (eye-level, low-angle worm’s-eye, bird’s-eye, drone, rooftop, etc.).
  • If there is a watermark, you must mention it.
  • Your response will be used by a text-to-image model, so avoid useless meta phrases like “This image shows…”, "You are looking at...", etc.
person_name STRING
max_new_tokens INT
temperature FLOAT
top_p FLOAT
top_k INT

Outputs

STRING

STRING

Extension: JoyCaption Nodes

Nodes for running the JoyCaption image captioner VLM.

Authored by fpgaminer

Looking for a different node?

Run ComfyUI workflows without the setup

No installs, no CUDA version roulette, no GPU sitting idle on your bill. Bring a workflow and run it in the browser.

Learn more