ComfyUI Node
Tile Test: Captions
A ComfyUI node in image/upscaling/tile testing with 18 inputs and 7 outputs.
Tile Test: Captions
- image
- layout
- clip
- captions
- tile_texts
- tiles
- prompt_tags
- tags_listed
- tags_verified
- tags_final
◄tiles►
◄with_neighborstrue►
◄position_termstrue►
◄caption_megapixels0.79►
◄tile_caption_max_tokens768►
◄global_style_max_tokens768►
◄tile_tags_verification_threshold1.0000►
◄tile_tags_position_threshold0.90►
◄prompt_tags_verification_threshold1.0000►
◄prompt—►
◄global_style_instruction—►
◄tile_caption_instruction—►
◄tile_tags_instruction—►
◄prompt_tags_instruction—►
◄tile_tags_verification_statement—►
Categoryimage/upscaling/tile testing
Inputs (18)
| Name | Type | Default | Description |
|---|---|---|---|
| image | IMAGE | The canvas from Tile Test: Upscale, at the layout's target size. | |
| layout | CATR_LAYOUT | The layout from Tile Test: Layout. The tiles are captioned from it. | |
| clip | CLIP | Must be a vision-language text encoder with a text generator (Krea 2 family). | |
| tiles | STRING | Comma separated tile numbers, as Tile Test: Layout labels them. Empty captions every tile. The same list comes out of the tiles output for Tile Test: Render. | |
| with_neighbors | BOOLEAN | true | Caption the bordering tiles of every named tile as well, which Tile Test: Render needs when its with_neighbors is on. Off captions the named tiles only. |
| position_terms | BOOLEAN | true | Score each kept tag on six strips of its tile, drop a tag no strip holds, and write a position term such as top-left. Off writes each tag without a term, makes no strip requests and drops no tag for its strips. Read by the tags kind when tile_tags_verification_statement is connected, and ignored by the caption kind. |
| caption_megapixels | FLOAT | 0.790–2 | How much of the picture the VL model reads for every caption this node writes, the tile captions and the style caption. Use 0 for the picture's own size, capped at 2.0 megapixels. The tags kind reads it for the style caption only, since its tags questions read a fixed copy of about 1 megapixel. Can take the caption_megapixels output of Tile Test: Settings. |
| tile_caption_max_tokens | INT | 7681–4096 | Generation budget for each tile caption, which also covers the model's hidden reasoning turn. Read by the caption kind and ignored by the tags kind. Can take the tile_caption_max_tokens output of Tile Test: Settings. |
| global_style_max_tokens | INT | 7681–4096 | Generation budget for the style caption, which also covers the model's hidden reasoning turn. Read when global_style_instruction is connected. Can take the global_style_max_tokens output of Tile Test: Settings. |
| tile_tags_verification_threshold | FLOAT | 1.00000.00001–1 | The score tile_tags_verification_statement must reach on the entire tile to keep a tag the tile's own list names. Read by the tags kind when tile_tags_verification_statement is connected. Can take the tile_tags_verification_threshold output of Tile Test: Settings. |
| tile_tags_position_threshold | FLOAT | 0.900–1 | The score tile_tags_verification_statement must reach on one of the six strips of a tile for that strip to hold a tag. A tag no strip holds is dropped, and a strip that is the only one holding a tag on its axis names the tag's position term. A center column adds no word beside a row word. Read by the tags kind when position_terms is on. Can take the tile_tags_position_threshold output of Tile Test: Settings. |
| prompt_tags_verification_threshold | FLOAT | 1.00000.00001–1 | The score tile_tags_verification_statement must reach on the entire tile to keep a thing from the prompt that the tile's own list lacks. Read by the tags kind when prompt_tags_instruction and tile_tags_verification_statement are connected. Can take the prompt_tags_verification_threshold output of Tile Test: Settings. |
| promptopt | STRING | Optional. The prompt the image was made from, as a text link. Connect the positive prompt's text. With the tags preset the VL model lists the things the prompt names, and a tile adds a listed thing to its tags when the VL model confirms it in that tile. In a caption preset it fills {PROMPT} in the instructions, and a preset with {PROMPT} needs it connected. The diffusion model never reads the prompt text itself. | |
| global_style_instructionopt | STRING | What the VL model is asked about the entire image, for one style caption placed on top of every tile's text, in either kind. {PROMPT} is filled from prompt. Unconnected, or holding only whitespace, writes no style caption. | |
| tile_caption_instructionopt | STRING | What the VL model is asked about each tile, which runs the caption kind. {PROMPT} is filled from prompt. Connect this or tile_tags_instruction, never both. Unconnected leaves the caption kind off. | |
| tile_tags_instructionopt | STRING | The question that asks the VL model to list the things in each tile as comma separated tags, which runs the tags kind. It cannot hold {PROMPT}. Connect this or tile_caption_instruction, never both. Unconnected leaves the tags kind off. | |
| prompt_tags_instructionopt | STRING | The question that asks the VL model to list the things the prompt names, once per picture. It must hold {PROMPT}, where the prompt goes. Every tile checks the listed things. Read by the tags kind when prompt is connected. Unconnected lists no things from the prompt. | |
| tile_tags_verification_statementopt | STRING | The statement each candidate tag is scored true or false against on its tile. It must hold {TAG}, where the tag goes, and a tag is kept at tile_tags_verification_threshold. Read by the tags kind only. Unconnected keeps every candidate unchecked and writes no position terms. |
Outputs (7)
| Name | Type | Description |
|---|---|---|
| captions | CATR_CAPTIONS | The style caption and every captioned tile's text, for the captions input of Tile Test: Render. |
| tile_texts | STRING | Markdown that lists the style caption once and then each captioned tile's own text, the texts Tile Test: Render conditions on. |
| tiles | STRING | The named tile numbers, comma separated, for the tiles input of Tile Test: Render. It is empty when the tiles input was empty. |
| prompt_tags | STRING | Markdown of the tags kind's prompt stage: the question sent once per picture, the VL model's reply and whether each listed tag was kept. |
| tags_listed | STRING | Markdown of the tags kind's listing stage: each tile's VL model reply and the tags parsed from it. |
| tags_verified | STRING | Markdown of the tags kind's verification stage: each tile's candidate tags with their scores and whether each was kept. |
| tags_final | STRING | Markdown of the tags kind's last stage: each tile's tags after the subset and position checks, and the tile text they make. |