CV EAST Text Detect
Find the text in a photo, boxes not letters
- image
- bboxes
- quads
- scores
- det_count
Scene text - the writing on a sign, a label, a shopfront, a licence plate, a UI screenshot - is a detection problem, not a segmentation problem. CV EAST Text Detect finds it and hands you boxes. It does not read the text; it tells you where the text is, which is usually the half you can't do with a threshold.
Why this one exists in this pack
The pack's README is unusually honest about its DNN nodes: everything goes through cv2.dnn on principle, and for several jobs (frame interpolation being the example they pick) core ComfyUI already does it better on the GPU in PyTorch. Text detection is not one of those jobs. There's no core EAST node, no core scene-text detector at all, so this is one of the genuine reasons to have the pack installed rather than a curiosity.
EAST itself is the Efficient and Accurate Scene Text detector - an old, small, still-surprisingly-good CNN. The node wraps it as cv2.dnn.TextDetectionModel_EAST, which does blob → forward → decode → NMS internally, and it runs in the pack's interruptible DNN worker (workers/dnn_tasks.py, task east_detect) so a long run can be cancelled instead of wedging the queue.
The network emits a per-pixel text score plus a geometry channel, at a quarter of the input resolution. Those geometry values are what give you rotated boxes - the reason EAST handles tilted signs that a plain rectangle detector mangles.
Inputs that matter
image- an IMAGE batch is processed frame by frame; a gray or BGR array is treated as one frame.model- a dropdown of the.onnxfiles you've put inComfyUI/models/onnx.conf_threshold(default 0.5) - the recall/precision dial. Lower finds more text and more texture noise.nms_threshold(default 0.4) - how aggressively overlapping boxes get merged. Raise it if you're getting the same line of text reported twice.input_size(default 320, range 320–1024, step 32) - the square the network works at. 320 matches the reference demo; bigger is more accurate and slower. If you're hunting small print, this is the knob.
Outputs and where they go
Four of them, and they're the reason this node is pleasant rather than merely useful:
bboxes- axis-aligned boxes as{x, y, width, height, score, label}dicts. That's ComfyUI'sBOUNDING_BOXtype, so the core Draw BBoxes node previews them with zero effort.quads-(N, 4, 2)float32 rotated corners, one per detection. Feed them toCV Draw Polygonfor a tight rotated outline; set its thickness to 0 for a filled polygon or a mask instead.scores-(N,)confidences, same order as the boxes. Empty when nothing is found.det_count- the total, for branching.
Zero detections is a valid result, not an error: empty arrays, count 0.
Getting the model
Models are not bundled with the pack - download east_text_detection_2026jul.onnx (Apache-2.0, from the OpenCV contribution collection on Hugging Face, opencv/opencv_contribution) and drop it in:
ComfyUI/models/onnx/east_text_detection_2026jul.onnx
The pack's own model_sources.txt records the URL under which the shipped example grants it, so check there before redistributing anything. Reload the ComfyUI page or restart after copying - the model combo is built from what's in the folder at load time, so a file you drop in mid-session won't appear until then.
Install
# ComfyUI Manager → search "ComfyUI CV" → install → restart
# or manually:
cd ComfyUI/custom_nodes
git clone https://github.com/bmad4ever/comfyui_cv
cd comfyui_cv && pip install -r requirements.txt
Python ≥ 3.12 and a recent ComfyUI (V3 node API) are hard requirements; the sole runtime dependency is opencv-contrib-python-headless~=5.0.0.93. Keep it contrib - a plain opencv-python install over the top wipes the contrib submodules, and the pack's tools/repair_opencv_contrib.py --check / --apply exists because there's no install-time guard. Everything here runs through OpenCV's DNN module, so it's CPU-side by default and won't compete with your sampler for VRAM in any clever way.
When it misbehaves
- Empty
modeldropdown - no.onnxinComfyUI/models/onnx, or you dropped it somewhere else. Also check you reloaded the page. - Small text missed - raise
input_sizetoward 640/1024. EAST's 1/4-resolution output is the limiting factor, not your threshold. - Everything looks like text -
conf_thresholdtoo low. Fine texture and repeating patterns are the usual false positives. - Nothing loadable - the README's general caveat applies: models converted to ONNX for this pipeline don't always load, and modern architectures are limited by both the DNN implementation and the pinned OpenCV version. If EAST's ONNX won't load, that's the class of problem, not a typo in your path.
- It's a detector, not OCR. You get geometry. If you need characters, that's a separate recognition step.
The author, bmad4ever, is a small-utility ComfyUI maintainer (he also ships a cartesian-product list node and an undo/redo extension) and this pack is his fork-turned-rewrite of opencv-comfyui, explicitly and loudly written with heavy LLM assistance and explicitly not production-grade. Use it to locate text; don't build a business on it.
Inputs (5)
| Name | Type | Default | Description |
|---|---|---|---|
| image | NPARRAY,IMAGE | Input image. An IMAGE batch is processed frame by frame; an NPARRAY (BGR/RGB/gray, any dtype) is treated as one frame. Accepts a ComfyUI IMAGE/MASK directly (frame 0 of a batch) or an NPARRAY. Arithmetic ops (add, multiply, etc.) process the full IMAGE batch when both inputs have the same batch size. | |
| model | COMBO | EAST .onnx model from ComfyUI/models/onnx. | |
| conf_threshold | FLOAT | 0.500–1 | Minimum confidence to keep a detection. Lower detects more (and more false positives). |
| nms_threshold | FLOAT | 0.400–1 | Non-maximum suppression IoU threshold for merging overlapping text boxes. |
| input_size | INT | 320320–1024 | Square input size for the network. Larger is more accurate but slower. 320 matches the reference demo. |
Outputs (4)
| Name | Type | Description |
|---|---|---|
| bboxes | BOUNDING_BOX | Axis-aligned {x, y, width, height, score, label} dicts - feed the core 'Draw BBoxes' node for a quick preview. |
| quads | NPARRAY | (N, 4, 2) float32 rotated corner points per detection. Feed into 'CV Draw Polygon' for precise rotated outlines (set thickness=0 for filled polygons or a mask). |
| scores | NPARRAY | (N,) float32 confidence per detection, same order as bboxes. Empty (0,) when nothing is detected. |
| det_count | INT | Total number of text regions detected. |