Nodes/FiL_Design_ImageMind/🕵️ Optic Scanner
ComfyUI Node

🕵️ Optic Scanner

Turn any image into a prompt your model will actually obey

By FiL-Design-Ai·Created 2 months ago·Updated 8 days ago· 8
🕵️ Optic Scanner
  • config
  • image
  • prompt
  • metadata_json
  • metadata_dict
agent⚪ None
agent_focus⚪ None
detail_levelnormal
languageen
model_typeAuto/None
video_duration0
video_aspectAuto
video_soundAuto
video_cameraAuto
prompt_modeAuto
photo_styleNone
nsfw_photo_styleNone
art_styleNone
nsfw_art_styleNone
seed-1
response_formattext
width0
height0
prompt
negative_prompt
custom_style

Optic Scanner is the node that matters most in FiL_Design_ImageMind - the README says so itself, and for once the author isn't overselling. Everything else in the LLM half either feeds this node or consumes what it writes. Feed it an image and it sends that image to a vision model and gets back a generation prompt tuned for a specific target: Z-Image, FLUX, SDXL, QWEN, Krea 2, Ideogram 4, or a universal video profile. Leave the image socket empty and it becomes a text-to-prompt expander instead. Both halves are genuinely useful, which is rare.

What the output is actually made of

The output is not one template with a text box. It's five independent axes that stack - agent, agent_focus, detail_level, model_type, and a style - plus your prompt and negative_prompt steering on top. The agent is the single biggest quality lever in the whole pack. Twelve subject lenses (Portrait, Products, Nature & Landscape, Fashion, Animals, Architecture, …) each tell the model exactly what fields to describe and what to ignore. Run a product photo through Products and you get "anodized unibody, MagSafe and two USB-C on the left edge." Run the same photo through Portrait and the model chases hair/skin/pose fields that aren't in the frame - output gets visibly worse. Match the agent to the subject.

agent_focus lays a craft layer on top without replacing the agent (📐 Composition, 💡 Lighting & Color, 🔬 Ultra Detail, 🎬 Cinematic, 🎭 Emotion & Motion). detail_level is a real word budget, not "more adjectives": tiny (20–50 words) up through ultra (500–1200). model_type rewrites phrasing and length ceiling for the target generator and never touches which facts got pulled out. response_format shapes the text - prose, flat comma tags, or JSON for the Ideogram/FLUX schema.

The two knobs that read backwards

negative_prompt never reaches the model as an SDXL-style negative. For FLUX/Z-Image/Krea 2/Ideogram 4/Video - none of whose APIs take a negative input - it's rewritten into positive constraints: "don't include blurry" becomes "express the scene without blurriness." Type the same words with SDXL and you get a plain Avoid: clause. And prompt, when an image is connected, is not a description - the image is the source of truth. It's a targeted instruction layered on top: "emphasise the fabric texture and stitching." It can't swap the subject (that comes from the image); it steers how the answer reads. Without an image, prompt becomes the entire idea to expand.

prompt_mode decides one call or two. Auto (default) runs a single Hybrid call until you pick a style, then switches to Two-Stage - stage one writes a plain factual description, stage two restyles that locked description. That's the more reliable path for style work, and it's why the style contracts can detect when a preset plainly doesn't match the photo and downgrade it to background flavour instead of inventing neon signs into a field. If a style "didn't take," check metadata_dict.response_outcome - it reports whether the final text actually used the required cues.

Wiring and debugging

Three outputs: prompt (into your CLIP Text Encode), metadata_json, and metadata_dict - a dict that carries sent_prompt, the exact system/user text of the LLM call that produced your result. When a generation comes out weird, that's where you look first. The width/height inputs are sockets, not panel fields - wire the target resolution in from Empty Latent Image and the prompt tailors its composition to that aspect ratio.

The biggest complaint people raise with LLM-in-graph prompts is subject drift and dirty output. This pack's answer is structure: the agent templates constrain what's described, and the metadata tells you what was actually sent. Two tight calls beat one freeform rewrite - the same lesson the community converged on for prompt-enhancer nodes generally.

Setup: it needs config from FiLProviderLoader. For first runs, pick Ollama as the provider (no key), match the agent to your subject, and leave model_type on Auto/None. If the model list is empty or vision calls fail, that's a provider problem, not a scanner problem - see the Provider Loader article.

Category🎨 FiL Design/🧠 LLM

Inputs (23)

NameTypeDefaultDescription
configFIL_PROVIDER_CONFIGSettings received from the Loader node.
agentCOMBO⚪ NoneSubject domain — what is in the frame.
agent_focusCOMBO⚪ NoneCraft layer to weigh heavier on top of the agent: composition, light, fine detail, cinematic read, or the subject's state and movement.
detail_levelCOMBOnormalHow detailed the answer should be.
languageCOMBOenLanguage of the answer. Auto replies in the language of your prompt.
model_typeCOMBOAuto/NoneTarget generation model — adjusts prompt syntax, length, and format. Video is a universal profile for video models (MiniMax H2/H3, Wan, HunyuanVideo, LTX Video, Kling); MiniMax H3 gets a dedicated timeline shot-block profile.
video_durationINT00–20Requested clip length in seconds. 0 = Auto (the LLM decides). The range follows the model: MiniMax H3 accepts 4-15 whole seconds (API limit), universal Video 2-20. Shown only for video model types.
video_aspectCOMBOAutoAspect ratio written into the shot framing — for MiniMax H3 into the timeline header line ('10s, 16:9.'). Auto leaves the choice to the LLM. Shown only for video model types.
video_soundCOMBOAutoSound design mode. Auto = the LLM decides; Off = silent clip (no sound clause); Layered = mandatory ambience + foley + music clause for audio-capable models (H3, Veo). Shown only for video model types.
video_cameraCOMBOAutoPreferred camera move. The LLM builds the shot around it and may adapt per story stage — a preference, not a hard lock. Shown only for video model types.
prompt_modeCOMBOAutoAuto picks Hybrid or Two-Stage depending on whether a style is selected.
photo_styleCOMBONonePhotographic style overlay applied on top of the base description.
nsfw_photo_styleCOMBONoneAdult-only photographic style overlay (alternative to Photo style).
art_styleCOMBONoneArt style overlay applied on top of the base description.
nsfw_art_styleCOMBONoneAdult-only art style overlay (alternative to Art style).
seedINT-1-1–999999999999Provider-side generation seed, if supported. -1 lets the provider pick one.
response_formatCOMBOtextOutput shape: prose text, flat comma-separated tags, or a JSON object.
imageoptIMAGEImage or batch to analyze. Leave disconnected to generate from Prompt field.
widthoptINT00–16384Target image width in pixels. Helps tailor prompt composition to aspect ratio if > 0.
heightoptINT00–16384Target image height in pixels. Helps tailor prompt composition to aspect ratio if > 0.
promptoptSTRINGWITHOUT AN IMAGE — the idea itself, expanded into a finished prompt: ginger cat on a windowsill, rain outside abandoned metro station, morning light through a hole in the ceiling WITH AN IMAGE — what to focus on and what the text is for: emphasise the light and skin texture prompt for an album cover the cup on the table is the subject, not the person keep empty space on the right for a headline WON'T WORK (this comes from the image): replace the woman with a man make her stand instead of sit remove the second person Leave empty and the image is simply described in detail. Details: docs/scanner-prompts.md
negative_promptoptSTRINGRemoves WORDS from the generated text, not objects from the image. Pose, subject and composition cannot be changed here — use the prompt field. PUT HERE what the model adds on its own: cliches: masterpiece, best quality, 8k, highly detailed effects: bokeh, lens flare, dramatic vignette junk: watermark, logo, signature, text overlay defects: extra fingers, plastic skin, waxy texture judgements: beautiful, stunning, breathtaking Short nouns, comma-separated, 5-10 of them. Not sentences with 'no': under FLUX / Z-Image Turbo / Krea 2 / Ideogram 4 / Video the list is flipped into positive wording, and 'do not make it blurry' has nothing to flip — write blurry. Not sure what to put? Leave it empty, run once, and pick out the words that bother you in the answer. Details: docs/scanner-prompts.md
custom_styleoptSTRINGCustom style text. Overrides photo/art/nsfw preset styles.

Outputs (3)

NameTypeDescription
promptSTRINGGenerated prompt text.
metadata_jsonSTRINGJSON metadata string.
metadata_dictDICTParsed metadata dict.