DOGMA ProposeCategories v1.0.11
Turning a VLM's rambling scene inventory into categories you can actually segment
- proposal
- presence_prompt
- report
DOGMA Nodes is not a general-purpose toolkit. It's the in-house pipeline of a commercial AI-video studio - the author is the Reddit user axior, part of the Dogma team that produced the VISA/Intesa Sanpaolo spot for Milano Cortina 2026 with "hundreds of VACE inpaintings in comfyui" and masks that, in his own words, "often had to be done manually frame by frame in after effects, since segmentation is not always working perfectly, especially when you point something very specific."
That sentence is the whole reason this pack exists. ProposeCategories v1.0.11 is the first half of the answer: it takes a vision-language model's description of your scene and turns it into a small list of machine-usable categories - each with a search phrase a text-prompted segmenter can act on.
How it works
You feed it planner_text, which in practice is the raw output of a planner node talking to a VLM. The node pulls the first JSON array out of that text and treats each object as one category record, then cleans up after the model the way you'd expect:
- The category name goes through the pack's canonicaliser, so
person,personsandpedestrianall collapse topeople- and duplicates are dropped. - Records with no evidence field, or a category in the pack's inactive set (
none,unused,n/a), get thrown away rather than passed downstream as a broken slot. - At most eight categories survive.
- Query phrases are filtered hard. This is the bit people miss: text SAM wants object nouns. The node maintains a small fixed vocabulary for common categories (
people→person,pedestrian;vegetation→tree,shrub) and bans scene-level filler -group,crowd,scene,background,city,italy, and anything containing1970orcinematic. If nothing survives it falls back to the category name itself.
Then it does something quietly smart. It builds presence_prompt, a second VLM instruction that only contains the proposed id and category - no evidence, no reasoning from the first model - and tells the model to independently check the actual photograph, answering present, absent or uncertain. You feed the first model's conclusions nowhere near the check.
The inputs and outputs that matter
Three things, exactly as the schema has them:
planner_text(STRING, required) - wire a VLM planner into it. It's a forced input, so you can't just type "a photo of a street" and expect categories.proposal(DOGMA_PLAN) - the bundle to send toDOGMA VerifiedPlan8 v1.0.11, which applies the presence verdicts.presence_prompt(STRING) - send this to your VLM, then to VerifiedPlan8 aspresence_text.report(STRING) - the parsed proposal as pretty JSON, for eyeballing without a text display node on the plan itself.
Install
Manager → search DOGMA Nodes, or:
cd ComfyUI/custom_nodes
git clone https://github.com/axior/ComfyUI-DOGMA-Nodes
Restart ComfyUI. The pack's only pip dependency is scipy>=1.10, but the nodes do look up other node classes by name at execution time - this one needs the planner upstream (DOGMABoundedPlannerV112 requires the comfyui_vlm_nodes pack for ModernVLM), and your ComfyUI needs to be recent enough to expose SAM 3 and lazy inputs.
Where people get burned
The parser here is deliberately dumb, and the changelog for 1.0.12 says why: on repetitive VLM output the inventory got truncated, the old code grabbed the last nested ] as the end of the array, and json.loads raised instead of coping. If you're on a chatty VLM that likes to keep listing individual people, v1.0.12's ProposeCategories is the same node with a real recovery path. Otherwise the usual hits are the boring ones: no usable categories (check the report), a category you know is in frame coming back absent from the presence check, and text SAM returning broad blobs because your query was a scene name rather than a noun. The presence check is probabilistic - the node's own preview text tells you it "can still be wrong; inspect the masks."
Inputs (1)
| Name | Type | Default | Description |
|---|---|---|---|
| planner_text | STRING | — |
Outputs (3)
| Name | Type | Description |
|---|---|---|
| proposal | DOGMA_PLAN | — |
| presence_prompt | STRING | — |
| report | STRING | — |