ComfyUI_toyxyz_test_nodes
This node was created to send a webcam to ComfyUI in real time. This node is recommended for use with LCM.
Nodes (5)
Your webcam, straight into the graph, no app required
Letterbox anything to a fixed canvas without stretching it
The brake pedal for real-time ComfyUI loops
The stable way to feed a live webcam into ComfyUI
Write your generations to any folder you want
ComfyUI_toyxyz_test_nodes
This is a custom node that collects the tools I use frequently.
https://github.com/toyxyz/ComfyUI_toyxyz_test_nodes/assets/8006000/8536e96a-514a-48b2-b1aa-8eccbd3fa853
(This video is at 4x speed)
Update
2026/09/29 Add booru tag prompter node
2026/09/16 Add image prompter node
2026/09/14 Add minimax h3 camera node
2026/08/25 Add Minimax-H3-prompter node
2026/04/18 Add Draw area mask, ComfyCouple Region multi, Crop area mask node
2026/04/16 - Add Anima support to ComfyCouple Region node
2025/10/30 - Add lora hook support to ComfyCouple Region node
2025/10/21 - Add Openpose Editor Node, Pose Interpolation, ComfyCouple Region, ComfyCouple Mask, Comfy Couple Region Extractor
2025/03/10 - Add Visual area mask node
2024/11/14 - Add Load Random Text From File node
2024/11/04 - Add Export glb node.
2024/11/02 - Add remove noise node for normal map. Added sobel ratio for more accurate Noraml.
2024/10/25 - Add depth to normal node.
2024/08/11 - Add Direct_screenCap node.
2023/11/24 - AddSave image to path node. Add Render preview, Add export video, Add face detection (After the update, you will need to run CaptrueCam/setup.bat one more time.)
2023/11/29 - Add Region Capture. Made the Webcam app UI smaller.
2023/12/01 - Add Ai render overlay
Installation
-
Git clone this repo to the ComfyUI/custom_nodes path.
git clone https://github.com/toyxyz/ComfyUI_toyxyz_test_nodes -
Run setup.bat in
ComfyUI/custom_nodes/ComfyUI_toyxyz_test_nodes/CaptureCam
Usage
Default workflow
(Workflow embedded)
Render preview workflow
(Workflow embedded)
Direct Webcam capture workflow (without webcam app)
(Workflow embedded)
Minimax-H3-prompter
Builds MiniMax H3 audiovisual prompts from a Shot/Move timeline and optional image, video, and audio references. It generates prompts, not the final AI video.
<img width="2618" height="1571" alt="image" src="https://github.com/user-attachments/assets/ceea6c19-1233-458e-a502-303a13a80115" />Quick start
- Choose a model, mode, and duration. Auto selects the mode from your references. Use 5–15 seconds as a practical starting range; the displayed frame count is H3-aligned.
- Enter the subject, action, setting, and any camera instructions in Prompt. + Shot starts a new take; + Move continues the same Shot without a cut. Drag timeline boundaries to adjust timing.
- Optionally set Visual style and Camera style for the selected Shot/Move. Visual style controls appearance; Camera style adds handling such as handheld or stabilized movement, without replacing the requested path or speed.
- Configure Camera, then select Generate Prompt. Enhance defaults to Normal: None is concise, Normal expands the request, and Strong produces a longer, richer description. Stop cancels generation.
- Connect
generated_promptandlengthto your H3 workflow. Enable Auto Run to generate the prompt when ComfyUI executes the node. Queue generation shows ComfyUI's native green progress bar by stage (not token count or remaining time). The separate Generate Prompt button keeps its own status log.
Camera and prompt display
Camera options are saved per Shot/Move: framing, viewpoint, lens height, angle, route, roll and composition. Camera level sets height independently of Angle. Movement range and Speed guide the prompt, not preview distance or speed. User camera instructions take priority; the proxy preview is illustrative, not a guarantee of the generated video's framing.
The display dropdown and Copy use the same text area:
- Generated prompt: the latest generated result.
- Raw prompt: system and user input; after generation, includes the actual reference analysis.
- Camera prompt: procedural camera timeline before Qwen applies user-text overrides.
References and models
Modes: T2VA (text), I2VA (first frame), FL2VA (first/last frames), L2VA (last frame), and REF2VA (reference generation/editing).
- Image: First frame, Last frame, Frame, or Subject. Drag Frame anchors to exact timeline positions; first/last anchors stay fixed. Use reference strength for subject retention.
- Video: select editing, continuation, or motion/action timing. Move and trim clips on the timeline; only the visible source interval is used. For motion transfer, identify source-to-target mappings in Prompt, e.g. “red figure = the woman; blue figure = the man.” Motion references include camera behavior but do not copy source appearance or scenery.
- Audio: select the intended reuse role. Loading a video alone does not request audio reuse.
Use aliases with @. Reference order must match downstream H3 slots.
Long reference lists, media timeline tracks, and Camera controls scroll inside the
node instead of expanding it. Saved node sizes remain adjustable; long prompt
text and logs do not set the minimum node height.
Outputs include image_N, frame_N, video_N, and audio_N as applicable.
Workflow restoration preserves existing output connections and saved prompts until
project and camera inputs are ready. If saved data or output names are ambiguous,
the execution log warns and preserves the existing slots instead of deleting them.
Reference videos are decoded sequentially into the selected 24 fps frames, preserving
source resolution, float32 output, display rotation, and trimmed audio. Large video
and camera buffers (256 MiB or more, or when RAM is low) use temporary file-backed
storage under test/cache/h3_video_memory/. Files are released when their last tensor
or cached output is released. Allow sufficient disk space; downstream resizing and
VAE encoding still need memory. Existing ComfyUI RAM-cache eviction is used without
unloading models or changing global settings.
Qwen3.8 27B supports all modes and Normal/Strong expansion. Each Qwen task reuses one llama-server for reference analysis and final writing, then releases it on completion, cancellation, or failure. Text-only tasks do not load the vision projector. Analysis and writing use separate requests, without accumulating the full image conversation in the final writing context. R2V includes each source analysis only once, even when multiple targets share it. Explicit source-to-target associations are carried across editing, continuation, and motion roles; these associations never override user instructions or change the selected role's scope. A motion target explicitly naming an existing image alias reuses that Subject instead of creating another person. Ambiguous mappings remain best-effort diagnostics, not reasons to block or regenerate the prompt. Compacted associations retain source-local actor IDs, selectors and target descriptions so action evidence remains linked to the correct target. Uncertain or mismatched appendices remain visible as unverified evidence. Model-inferred associations are input guidance only: output checks warn without adding correspondence sentences or overwriting motion-target definitions. They do not establish semantic correctness; the user's explicit instructions remain authoritative. MiniMax H3 Rewriter Omni supports all modes with its own expansion behavior. Missing models require download confirmation; a managed llama.cpp runtime is installed on first use. Qwen runs with a 16,384-token context shared by input and output: many timeline items or reference analyses can exceed it. Check the execution log for download progress, context warnings, and reference-mapping errors.
Camera sequence input
Connect minimax h3 camera through prompter_camera to use its duration and
enabled video/prompt guidance instead of the local Camera panel. See controls below.
Without this connection, Camera render optionally outputs the local panel's
preview sequence; it is not automatically added as a video reference.
minimax h3 camera
A 3D camera/keyframe editor for reference videos and procedural camera prompts.
<img width="2265" height="1498" alt="image" src="https://github.com/user-attachments/assets/07571bc7-ff03-46b8-af33-2d5e3c7dfa9e" />- Translate / Rotate / Scale, World / Local: edit the selected camera or subject. Middle-mouse drag pans the editor; Camera view toggles the output preview.
- Shape: Human, Box or Sphere; adjust subject color and transforms in the inspector. The white T on a Human's face marks its front.
- Free / Orbit: position the camera directly or move around a target. Aim selects Track target or Free rotation per key; Target height (0–1) selects the aim point from the subject's base to its top.
- + Key / Auto key: animate camera and subjects on separate tracks. Drag keys to retime; playback and Undo/Redo are available. Extra tracks scroll. Smooth / Linear / Hold applies from the previous key to the selected key. Hold keeps the previous pose, then jumps; a changed camera view creates a Shot cut.
- Duration (s), aspect ratio and MP: set length and render size. Output is 24 fps with H3-aligned frame counts; lower MP reduces rendering cost.
- Floor grid / Background grid: toggle the opaque floor/grid and spherical orientation grid in both previews and rendered video.
- refvid (default on): send rendered video to the connected prompter.
Choose its video role (motion/action, editing or continuation), alias and description there.
Turning it off removes that video reference and its
video_Noutput; reconnect downstream video links if re-enabled. - Use camera prompt (default on): send procedural camera motion and shot views to prompt generation. With refvid off, sends text only; both off sends neither. User instructions take priority. Continuation uses the source as history, not a route to replay.
Connect prompter_camera to the prompter's matching input. Its Cam Shot guide
shows camera cuts and frame ranges; playheads synchronize, while Shot/Move edits
remain separate. Generate a new prompt after changing the camera.
Standalone outputs: camera_render is the rendered IMAGE sequence;
camera_prompt is procedural text without Qwen. Both remain available regardless
of the two guidance toggles. Prompter text overrides do not change rendered geometry,
and generated-video camera accuracy is not guaranteed.
Cut Video
Trims a ComfyUI VIDEO with one signed frame count while keeping its embedded audio aligned.
frame_count: positive values keep that many frames from the beginning; negative values keep that many frames from the end (-1returns only the final frame and-22returns the final 22 frames);0keeps the complete connected mediainvert: when enabled, a positive value excludes that many frames from the beginning and a negative value excludes that many frames from the endvideo: required sourceVIDEO, including its embedded audio and frame-rate metadata- outputs: trimmed
video,images,audio, and the original inputfps, in that order
VIDEO FPS is preserved. The image and audio outputs are extracted from the same selected VIDEO
interval. If the absolute frame_count exceeds available frames, the complete source is returned
without padding.
Connect Video
Connects two compatible ComfyUI VIDEO inputs into one longer VIDEO. video_1 plays first and
video_2 follows it. Both videos must have matching FPS and frame dimensions. Embedded audio is
joined in the same order; when only one input contains audio, silence is inserted for the other
video so synchronization is preserved. Set smooth_transition above 0 to overlap that many ending
frames of video_1 with the opening frames of video_2. During the overlap, video and audio from
video_1 fade from 100% to 0% while video_2 fades in. A value of 0 performs a direct join.
image prompter
<img width="1995" height="1461" alt="image" src="https://github.com/user-attachments/assets/48f93f03-23c9-4371-9d16-aefce5c35c08" />Turns requests into Anima tag-and-caption prompts, English scene prompts, or image-editing instructions using local Qwen. Optionally connect reference images and an image prompter preset node.
- Enter your request in
prompt. - Optionally connect
image_1and/orpreset. Connecting an image reveals the next input, up toimage_10. - Run the workflow, then use the
promptoutput with a text encoder or text display.
Controls
llm_model: shows the supported Qwen model and installation status.seed: controls prompt variation, not the image generator's seed. Usefixedfor repeatable tests.prompt_type:animawrites a tag-forward hybrid prompt: validated Danbooru tags followed by concise English sentences covering every user-specified placement, pose, limb position, gaze, expression, action, and object location. Complex inputs use as many brief sentences as their explicit spatial facts require. It accepts tags alone, natural language alone, or both.defaultuses medium-first prose descriptions, explicit spatial relationships, and user-information preservation (formerlynormal_2). Oldernormalandnormal_2selections migrate todefault. Enhance strength is a separate setting.qwen_image_2.1: writes editing instructions, identifying what to change and preserve. Uses the existing Qwen writer, not a separate official Prompt Enhancer model. Without images, it rewrites the editing request from text only.enhance:nonetranslates and organizes;normaladds detail;strongdevelops open details more richly. Every level preserves explicit user information rather than summarizing it. No target word count is imposed; check generated text for model omissions.
In anima, recognized input tags are mapped to the bundled 2026-09-24 Danbooru vocabulary. Generated candidates are looked up by exact name, bundled alias, a few explicit semantic equivalents, then conservative spelling similarity and component lookup. Ambiguous or unsupported candidates are dropped rather than mapped to unrelated tags. rapidfuzz, when available, enables the spelling step; exact, alias, and component lookup work without it. Unknown user-authored tag tokens are retained. General tags use lowercase and spaces, score tags keep underscores, and recognized user-authored artist tags receive one leading @. No quality, safety, score, artist, or style tag is added by default. The short scene line describes only spatial relations and actions; appearance, clothing, light effects, style, and quality belong in tags. The dictionary snapshot and attribution are in nodes/data/README.md.
For Anima gaze and perspective, looking at camera maps to looking at viewer, while facing camera maps to facing viewer without asserting eye contact. Viewing position uses tags such as from behind and from below; scene prose uses viewer or viewpoint for these relations. A physical camera explicitly held or placed in the scene remains a camera object and can produce a camera tag.
In initial image prompter generation, explicit numeric ComfyUI weights such as (from front:4.92) retain their exact text and number. This applies to Anima and the prose prompt types even when Qwen omits or changes a weighted term. Prompt edits can intentionally change or remove weights, so the original weights are not restored after an Edit Prompt action.
For anima, enhance primarily controls tag expansion. none maps supplied facts; with tags alone, it returns only those recognized/custom tags and does not invent a scene line. normal asks for about 6–10 compatible optional tag candidates, and strong asks for about 16–24 plus a second pass for about 8–12 more when the user establishes a setting. Without an authored setting, the strong second pass adds only compatible lighting effects; clothing or sunlight does not establish a beach, sky, or other location. These are candidate targets, not forced output counts: dictionary validation and source fidelity can reduce the final count. For person appearance, clothing, and gaze extracted from natural-language input, Qwen must quote the exact source phrase; unsupported details are removed. The second strong pass cannot add subject appearance or clothing details and also checks the scene sentences against the user's pose and spatial instructions. Explicit user facts and limits remain the priority. Appearance leakage in the scene is repaired or removed locally while valid pose sentences are retained.
If the Anima writer returns empty or unusable output, or reaches its output token limit twice, generation logs a warning and emits the user's recognized/custom tags plus any authored natural-language text. This fallback preserves the source but may leave its language untranslated and cannot provide the requested enhancement. The other prompt types retain their existing error handling.
Anima's main writer, bounded retry, tag enrichment, and scene repair each allow up to 8,192 output tokens. The local Qwen server still has a 16,384-token context shared by input and output; a response can end sooner when that context is exhausted.
Qwen analyzes connected references together before writing the prompt. Specify source roles with <image1> through <image10>, for example: “Use <image1> as the canvas; replace only its bag with the bag from <image2>.” Only the first batch frame per socket is used. Existing image connections migrate to image_1.
Use the numbered image_1–image_10 outputs to pass the original images to the downstream editor in matching reference order. These outputs preserve the original tensors, resolution and full batches, not the resized first-frame analysis copies; an unconnected input returns no image. The prompt output remains first. Disconnecting an interior input does not renumber later references: avoid gaps when using a downstream editor that numbers references consecutively. “Only change…” and “keep… unchanged” take priority over presets and enhancement.
Editing the output
The generated prompt appears at the bottom of the node.
- Edit Prompt: enter an instruction and click OK to revise the current output with Qwen.
- For
defaultandqwen_image_2.1, editing first interprets the requested change, then revises the complete prompt. Foranima, editing revises the complete tag list and short scene line while preserving unrelated custom tags. An unchanged response is reported in the dialog instead of being applied as a successful edit. - Undo: restore the prompt before the last edit.
- Regenerate from inputs: clear the edited output, then run the workflow to generate from the inputs again.
Changing prompt_type clears the edited output, displayed prompt and Undo state; run the workflow to generate in the new mode. Other input changes leave the edited override active until you use Regenerate from inputs. Loading a saved workflow preserves its saved edit. These buttons do not start image generation. Unchanged inputs may use cached results; change the seed for a new variation.
Setup: uses the shared H3 Qwen/llama.cpp runtime. Existing weights are reused; missing model weights download on first use (about 16.8 GB for the language model, plus vision weights when needed).
Booru tag prompter
<img width="2440" height="1667" alt="image" src="https://github.com/user-attachments/assets/632856ae-1dbb-45c0-811d-c908da153d0b" />Enter tags in tags and connect the tags output to a text encoder. Autocomplete uses the bundled Danbooru tag list: type a fragment, then use ↑/↓ and Enter or Tab, or click a result. Selecting shiroko_(blue_archive), for example, inserts shiroko \(blue archive\). Manually typed text is not rewritten.
| Menu / input | Function |
| --- | --- |
| W (Wildcards) | Lists files in wildcards/. Click a file to insert __name__; each execution replaces it with one random non-empty line from that file. Folder opens the folder and ↻ refreshes the list. |
| F (Favorites) | Save stores the entire current prompt. Click an entry to insert it at the cursor; right-click to edit or delete it. ↻ refreshes the list. |
| Wiki | On first use, asks to download the offline database from Hugging Face (about 272 MiB). Later uses work offline. Browse categories, search, follow wiki links, and use Insert on a tag entry to add it at the cursor. Drag the title bar to move the window or its bottom-right corner to resize it. |
| camera | Optionally connect booru tag camera; its selected camera guidance is appended after your text. |
The node does not use Qwen or invent additional tags. Wildcard file contents and manually entered prompt text are preserved as written.
Booru tag camera
Connect its camera output to booru tag prompter.camera.
| Menu | Function | | --- | --- | | 3D preview | Shows an illustrative camera position. Drag a colored ring to change the horizontal or vertical view; release to snap to a preset. Scroll to change framing. | | Vertical view | Selects a view from above or below. | | Horizontal view | Selects front, back, side, left/right side, or a front/rear 45° view. | | Framing | Sets subject coverage, from close-up to very wide shot. | | Angle | Rotates the image view (Dutch, sideways, or upside-down). | | Perspective / Depth | Adds one perspective or projection tag, such as fisheye or isometric. | | Focus / Blur | Combines toggleable effects such as depth of field, bokeh, lens flare, and motion blur. Random samples a combination; Strength applies to the whole group. |
Each main list also has Random, which selects a non-None option on each run. Weights range from 0–10 (default 2.0): drag the number to adjust it or click to type; the reset icon restores 2.0. The preview is a guide, not a guarantee of the generated angle. Camera options do not choose a pose, outfit, background, or number of people.
image prompter preset
Supplies optional shot, angle, and style guidance to image prompter. Connect its preset output to the prompter's preset input. Use preset_prompt to inspect the preset text.
| Control | Purpose |
| --- | --- |
| shot_size | Extreme wide, wide, full body, cowboy, medium, medium close-up, close-up, or extreme close-up. |
| angle | Front/side/rear, high/low, overhead, drone/aerial, Dutch, over-the-shoulder, POV, or isometric views. |
| style_category | Filter styles by category; All shows all 115. This filter adds no prompt instructions. |
| style | Choose a rendering style. Categories cover photorealistic and cinematic looks, illustration and painting, 3D/craft/print, film grading, smartphone/social and editorial photography, director-inspired cinema, 2D animation, and amateur photo imperfections. |
None leaves that setting to the user prompt. To use style alone, set both shot_size and angle to None. Switching categories resets an incompatible style to None.
Explicit user instructions override presets. Camera settings guide framing, not the subject's pose, clothing, or background. Selected styles guide the whole image unless the user specifies otherwise. Exact framing and style fidelity depend on the image model.
Visual area mask
Creates masks for the specified regions. Useful for regional prompting.
Image_width: Specify the width of the mask
Image_height: Specify the height of the mask
area_number: Specify the number of areas to create. Maximum 12.
area_id : Area number to adjust. Starts from 0.
x : X position of the area selected in area_id.
y : Y position of the area selected in area_id.
width : Width of the area selected in area_id.
height: Height of the selected area at area_id.
strength: Strength of the selected area at area_id.
mask_overlap_method: default, subtract - Subtracts the masks from other regions from a single mask.
Update outputs: Update nodes according to the number in area_number.
<img width="1576" height="1591" alt="image" src="https://github.com/user-attachments/assets/dcc54f06-7d5c-4a2c-844c-11f8ec8088ae" />Draw area mask
Create a mask for regional prompting. Use Ctrl + click to select an area, and Alt + click to remove it.
<img width="1697" height="1526" alt="image" src="https://github.com/user-attachments/assets/7c570775-5ae8-4eea-bb56-67164a0c1f69" />Openpose Editor Node
Modify each body part of OpenPose
<img width="2140" height="1800" alt="image" src="https://github.com/user-attachments/assets/c54a6c62-a8aa-4418-a8b7-6bd65d5cce82" />Pose Interpolation
Generate interpolated poses between two OpenPose poses.
<img width="2408" height="1558" alt="image" src="https://github.com/user-attachments/assets/c76ae523-d09b-4d2a-b21f-447c76fdf36e" />ComfyCouple Region / ComfyCouple Mask
Regional Prompting Node. Supported models are SD 1.5, SDXL, and Flux, Anima. To disable Auto_inject_flux, you must free the model cache. To use Lora_hook, set skip_positive_conditioning to false. If you connect the ComfyCouple Base Prompt and ComfyCouple Background Prompt to the ComfyCouple Region, they will function as the base prompt and the background prompt.
<img width="3442" height="1324" alt="image" src="https://github.com/user-attachments/assets/1186f43e-4599-4aaa-8f89-6a38eac1fcc1" />Comfy Couple Region Extractor
Cut out the masked region from the couple region. It can be used in the face detailing workflow.
<img width="2505" height="1667" alt="image" src="https://github.com/user-attachments/assets/5b8871f0-24db-46e8-b69a-4c0e8aa844cf" />Booru tag prompter wildcards
Add UTF-8 text files to this custom node's wildcards/ folder. Each non-empty
line is one random option; it can contain multiple words, sentences, tags or weights.
Use __cloth__ for cloth.txt or __outfits/cloth__ for a subfolder file.
(__cloth__:1.5) preserves the weight around the chosen line. Nested calls are
supported with cycle/depth protection. Missing or unreadable calls remain literal.
Wildcards are expanded before camera composition without reformatting their text;
authored tag order is retained. Every occurrence samples independently on each execution
(consecutive runs may coincidentally select the same line).
Type __ in the prompt field to search the wildcard files inline. Use Up/Down,
Enter, Tab, Escape, or click just like tag suggestions. The right-side Wildcards
tab also lists files in a scrollable panel. Click a file to insert its call at
the cursor or replace selected text. Refresh updates the list
after filesystem changes. Calls are highlighted in the editor; expanded output
does not overwrite the original editable input. The screenshot's brace/multi-select
syntax is not part of this file-based wildcard implementation.
Load Random Text From File
Retrieves the entire text or random lines from a txt file at the entered path.
file_paht : The path to the text file or the path where the files are located
seed : seed for random line
edit_text : Edit tag_(tag) to tag (tag)
get_random_line : Get random line from txt. False for get entire text
get_random_txt_from_path : Randomly use one of all text files located in the entered path instead of one text file.
strength : Adjust the strength of the prompt.
ban_tag : Prompts to exclude.
text : Multi-line text as an alternative to text files
use_index : Gets text from the line corresponding to index instead of a random line.
index : The index of the line.
Export glb
Export a flat .glb file with a color image, normal map, and alpha mask.
You can specify the roughness, metallic, and save path.
Remove noise
guided_first : Apply guided filter first.
Remove noise from an image. Can be used to clean a normal map.
bilateral_loop: The number of times to apply the bilateralFilter. If 0, it is not used.
d/sigma_color/sigma_space : bilateralFilter parameters
guided_loop: The number of iterations of the guidedFilter. If 0, it is not used.
radius/eps: guidedFilter parameters.
Depth to normal
Converts a depth image to a normal map. It works very well with 2D images and DepthAnything v2.
depth_min : Depths lower than this value are replaced with 0.
blue_depth : Adjusts the intensity of the blue channel of the normal map to emphasize depth. The lower this number, the stronger the depth.
sobel_ratio : Makes the Normal map more stereoscopically accurate. Values between 0.1 and 0.3 are recommended.
Direct_screenCap
Captures an image from a specified window or screen.
capture_mode Default : Capture a defined area of the monitor window : Captures the area of the window entered in target_window window_crop : Same as window, but cuts off and captures the area relative to that window.
target_window : Name of the window to capture. You can find its name in the list of windows in the Webcam app.
Load Webcam Image
Load an image from a path.
To use this node with webcam, you must first run run.bat in ComfyUI/custom_nodes/ComfyUI_toyxyz_test_nodes/CaptureCam.
And in the Webcam app, you'll need to select your webcam and run capture with start.
Capture Webcam
Captures an image directly from the webcam selected with 'select_webcam'. (Usually 0)
This is unstable compared to the Load Webcam Image node.
If you're using obs, I recommend using the Load Webcam Image node.
Save image to path
This node saves the generated images to a defined folder path.
Set name to choose the base file name. In the Save method, images are saved without overwriting existing files by appending a number such as comfyui_000001.png. In Overwrite mode, an existing image with the same name is replaced.
Connect the MASK output from ComfyUI's Load Image node to preserve transparency when saving RGBA PNG images.
Connect or enter optional text to save a .txt file with the same name and folder as the saved image.
Load image from path
This node loads an image from a folder path by zero-based index. If index is larger than the number of images, it loads the last image. It outputs the image, mask, and selected image file name without extension.
LatentDelay
Set the delay between image generation.
ImageResize_Padding
Resizes the image while maintaining its proportions and painting the margins with the color you specify.
Webcam app
This script captures the selected webcam and saves it as an image file in real-time.
You can specify the resolution, format, and path of the image to be saved.
If you don't enter a path, it will be saved to the default path.
You can combine a sequence of saved images into a video using the Export button. Set the Save Image to Path save method to Save so each frame gets a unique numbered file name.
To use AI Render, you need a Save Image to Path node.
If you entered a location other than the default path in Save Image to Path, you must select a newly created render image in Select Rendered image.
Run_hide_cmd.vbs : Hide the cmd and run the app.
Webcam : List of camera devices connected to your computer
Width: Sets the width of the image.
Height: Sets the height of the image.
If either the width or height is zero, it will be automatically adjusted to fit the other values entered.
FPS : Set how often the capture occurs. If you enter 0, it is unlimited.
Webcam(Checkbox) : Preview the captured image.
Al Render: Preview the generated image in ComfyUi.
Always on top : Webcam, AI Render is always visible on top. To disable it, you need to close the preview window.
Face detect : Automatically recognize faces and generate masks. It is stored as face_mask.jpg. Use with inpainting.
Keep aspect ratio : Correct the aspect ratio of the image and the capture window.
Capture Path: The path where the captured image will be saved.
Render image: Path to the image generated by ComfyUI. Required to use AI Render.
If you don't enter a path, the default path is used.
Save format: Set the image format to be saved
Overlay alhpa : The alpha value of the overlay image displayed above the region capture window.
Padding: Set how to fill the margins of the image when using Keep aspect ratio.
Export video: Combines the image sequences located in the render image folder into a single video. Enter the desired FPS value.
Clear after export: Deletes the image sequence after the video is exported.
Add region window: Creates a window to specify the region to capture.
The name of the window is the same as entered in Save name. If you enter comma-separated text (e.g., A,B,C), you can use one Region window to capture three images, A, B, and C, alternating between them.
Save name: Set a name for the captured image.
Window list: Select the windows to capture. Other windows will not be captured. If set to Disable, it will be captured as it is displayed on the screen. Window capture is often unstable depending on the program. Be careful when using it.
Reload list: Refresh the list of webcams and windows.
The name of the currently captured and saved image file and the selected Region window are displayed in the Webcam/Ai Render preview window.
If you select 'Region Capture' from the Webcam list, it will capture the region of the window added with 'Add Region window'. If you select 'Window Capture', it will capture the entire window selected in the window list.
https://github.com/toyxyz/ComfyUI_toyxyz_test_nodes/assets/8006000/8723e014-caa5-4e16-8c8c-5c5edac6f141
Hotkeys :
S: Saves the image displayed in the current AI render to the render folder. Activate (click) the AI render window and then use it.
F : Toggles the Region widow (target window list selected) to full window mode. Select the target Region window and then click Use.
A : Select the Region window currently being captured. Enable the AI render or Webcam preview window, then use it.
C: Change the region window. Select the desired Region window and press the key to switch to that window. When used in the Ai Render or Webcam window, the Region windows are switched in the order they were created.
X: If the Region windows were created separated by commas, pressing the key while Ai Render or Webcam is active will switch the name of the captured image.
Q: Copy the image displayed in AI render to the clipboard. Activate AI render and then use.
M: Mask paint toggle. Allows you to paint the mask directly in the Webcam preview window. Paint with the left mouse and erase with the right. You can change the brush size with the mouse wheel. The mask is saved as the captured image name + '_mask'.
N: Erase all painted masks.
Z: Pause image capture. Use for Webcam or AI Render.
P: Displays the AI Render image as an overlay on the currently selected Region window. This is unstable and should be used with caution. It can only be enabled/disabled when capture is stopped and requires a target window to be set. Does not run even if there is no render image to load.
Render preview
Load the image file saved with the Save image to path node. Pressing 'Q' while the window is active will copy the preview image to the clipboard.
Face detection
Detect faces and create masks. Use it for inpainting with the Load Webcam Image node.
Note
The ControlNet preprocessor slows down the process, so I recommend using other tools to prepare the ControlNet image.
If you want ComfyUI to run continuously, use Auto Queue.
For maximum speed, set the VAE to taesd.