ComfyUI Node

Tile Test: Captions

A ComfyUI node in image/upscaling/tile testing with 18 inputs and 7 outputs.

By Blakeem·Created 2 months ago·Updated 3 days ago· 25
Tile Test: Captions
  • image
  • layout
  • clip
  • captions
  • tile_texts
  • tiles
  • prompt_tags
  • tags_listed
  • tags_verified
  • tags_final
◄tiles►
◄with_neighborstrue►
◄position_termstrue►
◄caption_megapixels0.79►
◄tile_caption_max_tokens768►
◄global_style_max_tokens768►
◄tile_tags_verification_threshold1.0000►
◄tile_tags_position_threshold0.90►
◄prompt_tags_verification_threshold1.0000►
◄prompt—►
◄global_style_instruction—►
◄tile_caption_instruction—►
◄tile_tags_instruction—►
◄prompt_tags_instruction—►
◄tile_tags_verification_statement—►
Categoryimage/upscaling/tile testing

Inputs (18)

NameTypeDefaultDescription
imageIMAGEThe canvas from Tile Test: Upscale, at the layout's target size.
layoutCATR_LAYOUTThe layout from Tile Test: Layout. The tiles are captioned from it.
clipCLIPMust be a vision-language text encoder with a text generator (Krea 2 family).
tilesSTRINGComma separated tile numbers, as Tile Test: Layout labels them. Empty captions every tile. The same list comes out of the tiles output for Tile Test: Render.
with_neighborsBOOLEANtrueCaption the bordering tiles of every named tile as well, which Tile Test: Render needs when its with_neighbors is on. Off captions the named tiles only.
position_termsBOOLEANtrueScore each kept tag on six strips of its tile, drop a tag no strip holds, and write a position term such as top-left. Off writes each tag without a term, makes no strip requests and drops no tag for its strips. Read by the tags kind when tile_tags_verification_statement is connected, and ignored by the caption kind.
caption_megapixelsFLOAT0.790–2How much of the picture the VL model reads for every caption this node writes, the tile captions and the style caption. Use 0 for the picture's own size, capped at 2.0 megapixels. The tags kind reads it for the style caption only, since its tags questions read a fixed copy of about 1 megapixel. Can take the caption_megapixels output of Tile Test: Settings.
tile_caption_max_tokensINT7681–4096Generation budget for each tile caption, which also covers the model's hidden reasoning turn. Read by the caption kind and ignored by the tags kind. Can take the tile_caption_max_tokens output of Tile Test: Settings.
global_style_max_tokensINT7681–4096Generation budget for the style caption, which also covers the model's hidden reasoning turn. Read when global_style_instruction is connected. Can take the global_style_max_tokens output of Tile Test: Settings.
tile_tags_verification_thresholdFLOAT1.00000.00001–1The score tile_tags_verification_statement must reach on the entire tile to keep a tag the tile's own list names. Read by the tags kind when tile_tags_verification_statement is connected. Can take the tile_tags_verification_threshold output of Tile Test: Settings.
tile_tags_position_thresholdFLOAT0.900–1The score tile_tags_verification_statement must reach on one of the six strips of a tile for that strip to hold a tag. A tag no strip holds is dropped, and a strip that is the only one holding a tag on its axis names the tag's position term. A center column adds no word beside a row word. Read by the tags kind when position_terms is on. Can take the tile_tags_position_threshold output of Tile Test: Settings.
prompt_tags_verification_thresholdFLOAT1.00000.00001–1The score tile_tags_verification_statement must reach on the entire tile to keep a thing from the prompt that the tile's own list lacks. Read by the tags kind when prompt_tags_instruction and tile_tags_verification_statement are connected. Can take the prompt_tags_verification_threshold output of Tile Test: Settings.
promptoptSTRINGOptional. The prompt the image was made from, as a text link. Connect the positive prompt's text. With the tags preset the VL model lists the things the prompt names, and a tile adds a listed thing to its tags when the VL model confirms it in that tile. In a caption preset it fills {PROMPT} in the instructions, and a preset with {PROMPT} needs it connected. The diffusion model never reads the prompt text itself.
global_style_instructionoptSTRINGWhat the VL model is asked about the entire image, for one style caption placed on top of every tile's text, in either kind. {PROMPT} is filled from prompt. Unconnected, or holding only whitespace, writes no style caption.
tile_caption_instructionoptSTRINGWhat the VL model is asked about each tile, which runs the caption kind. {PROMPT} is filled from prompt. Connect this or tile_tags_instruction, never both. Unconnected leaves the caption kind off.
tile_tags_instructionoptSTRINGThe question that asks the VL model to list the things in each tile as comma separated tags, which runs the tags kind. It cannot hold {PROMPT}. Connect this or tile_caption_instruction, never both. Unconnected leaves the tags kind off.
prompt_tags_instructionoptSTRINGThe question that asks the VL model to list the things the prompt names, once per picture. It must hold {PROMPT}, where the prompt goes. Every tile checks the listed things. Read by the tags kind when prompt is connected. Unconnected lists no things from the prompt.
tile_tags_verification_statementoptSTRINGThe statement each candidate tag is scored true or false against on its tile. It must hold {TAG}, where the tag goes, and a tag is kept at tile_tags_verification_threshold. Read by the tags kind only. Unconnected keeps every candidate unchecked and writes no position terms.

Outputs (7)

NameTypeDescription
captionsCATR_CAPTIONSThe style caption and every captioned tile's text, for the captions input of Tile Test: Render.
tile_textsSTRINGMarkdown that lists the style caption once and then each captioned tile's own text, the texts Tile Test: Render conditions on.
tilesSTRINGThe named tile numbers, comma separated, for the tiles input of Tile Test: Render. It is empty when the tiles input was empty.
prompt_tagsSTRINGMarkdown of the tags kind's prompt stage: the question sent once per picture, the VL model's reply and whether each listed tag was kept.
tags_listedSTRINGMarkdown of the tags kind's listing stage: each tile's VL model reply and the tags parsed from it.
tags_verifiedSTRINGMarkdown of the tags kind's verification stage: each tile's candidate tags with their scores and whether each was kept.
tags_finalSTRINGMarkdown of the tags kind's last stage: each tile's tags after the subset and position checks, and the tile text they make.