Nodes/ComfyUI-MiniMaxH3-Studio/H3 Reliability Status
ComfyUI Node

H3 Reliability Status

H3 Reliability Status reads the retry verdict — it never presses the button

By rookiestar28·Created 2 months ago·Updated a day ago· 79
H3 Reliability Status
    • status
    • progress
    • recovery
    • disclosure
    ◄operation_idh3.context►
    ◄stageplanning►
    ◄completed_units0►
    ◄total_units1►
    ◄run_staterunning►
    ◄cancel_requestedfalse►
    ◄attempt1►
    ◄max_attempts2►
    ◄checkpoint_statusincompatible►
    ◄resume_requestedfalse►

    The first thing to be clear about: this node does not retry anything. It doesn't start a worker, resume a checkpoint, or honour a cancellation by itself. Its own README line is that it "never starts, retries, resumes, or cancels a worker implicitly," and the code matches - it takes your declared state and computes a verdict.

    So what's the point? Long video jobs are where ComfyUI users actually suffer: a 12-second H3 segment dies at 80%, and you have no idea whether the checkpoint you kept is still usable or whether you're about to burn twenty minutes on a rerun that will fail identically. H3 Reliability Status exists to make that judgement explicit and typed, instead of a shrug and a queue button.

    The mechanism: a decision function with a witness

    You declare ten things. All are required:

    • operation_id (h3.context by default) - a bounded identity string, no paths or credentials.
    • stage - extraction, provider, planning, rendering, validation, host. Default planning.
    • completed_units / total_units - integers, 0 and 1 by default, up to a million each.
    • run_state - pending, running, completed, cancel_requested, cancelled, retryable_failure, failed, stale, recovered. Default running.
    • cancel_requested, resume_requested - booleans.
    • attempt / max_attempts - 1 and 2 by default, both capped at 8.
    • checkpoint_status - none, compatible, stale, incompatible. Default incompatible.

    From that it produces four outputs: status (H3_EXECUTION_STATUS), progress (H3_PROGRESS_EVENT), recovery (H3_RECOVERY_DECISION), and disclosure (a STRING for a human). The decision is one of a closed set: start, retry, resume, complete, cancelled, reject_stale, reject_incompatible, terminal_failure. There's also an execution_allowed flag on the status.

    The rules are the interesting part, and they're unusually strict in the right direction:

    • If cancel_requested is true you get cancelled and execution_allowed: false. Cancellation means no successful artifact is allowed - it doesn't mean "let the last step finish."
    • A stale checkpoint gets reject_stale ("stale checkpoint requires explicit restart"). An incompatible one gets reject_incompatible, full stop.
    • retryable_failure only earns a retry while attempt < max_attempts; past that it's a terminal failure.
    • Declaring attempt above max_attempts, or either above 8, is rejected outright - no unbounded retry loops encoded as data.

    The trap in the defaults

    checkpoint_status defaults to incompatible. If you leave it alone and ask to resume, the node correctly refuses - it cannot know your checkpoint matches the run, so it assumes it doesn't. That's fail-closed behaviour and it's the right default, but it's also the thing that makes people think the node is broken. If you genuinely have a compatible checkpoint, you have to say so, and it should be a fact you know, not a guess.

    Second trap: attempt > max_attempts is a validation error, not a forced clamp. If you're driving this from a prompt-expression or a script, keep the two numbers consistent.

    Where it goes in a graph

    Honestly? Often nowhere. Nothing else in the pack consumes H3_EXECUTION_STATUS, and the bundled example workflow (m7_05_h3_context_reliability) places it beside a Request → Plan → Compiler → Validator → Preview chain without wiring it into it. Treat it as a typed dashboard: it documents an orchestration decision - yours, the sidebar's, or your own automation layer's - in a form that a report, a log or a host UI can read without parsing prose. If you're building a wrapper around ComfyUI's queue and want "did this run get to retry or not" as structured data instead of a guess, this is the node.

    Because the outputs carry no prompt text, no paths and no credentials, they're also safe to log or screenshot, which is the same posture the API/wrapper-node writing argues for: keep the boundary visible, keep secrets out of the artifact.

    Install it

    The pack isn't in the Comfy Registry yet, per its README, so clone it:

    cd ComfyUI/custom_nodes
    git clone https://github.com/rookiestar28/ComfyUI-MiniMaxH3-Studio.git
    

    Restart ComfyUI and it's under Add Node → h3_context → runtime. No Python dependencies, no model downloads, no H3 weights needed - it's pure bookkeeping, so it runs fine while you're still setting up the rest. Same author as the ComfyUI-Doctor diagnostics pack, and the house style shows: fail-closed defaults, bounded inputs, and a refusal that tells you which field caused it.

    Categoryh3_context/runtime

    Inputs (10)

    NameTypeDefaultDescription
    operation_idSTRINGh3.contextBounded operation identity; no paths or credentials.
    stageCOMBOplanningExplicit long-running stage.
    completed_unitsINT00–1000000—
    total_unitsINT11–1000000—
    run_stateCOMBOrunning9 options: pending, running, completed, cancel_requested, cancelled, retryable_failure, +3
    cancel_requestedBOOLEANfalseExplicit caller cancellation request.
    attemptINT11–8—
    max_attemptsINT21–8—
    checkpoint_statusCOMBOincompatibleExplicit caller checkpoint compatibility result.
    resume_requestedBOOLEANfalseExplicit caller resume request.

    Outputs (4)

    NameTypeDescription
    statusH3_EXECUTION_STATUS—
    progressH3_PROGRESS_EVENT—
    recoveryH3_RECOVERY_DECISION—
    disclosureSTRING—