ComfyUI Node

react-1

Re-animate a face to a new performance — the deepfake-adjacent node

By Runware·Created 2 years ago·Updated 17 days ago· 140
react-1
    • video
    video
    audio
    numberResults1
    providerSettings.sync.activeSpeakerDetectionfalse
    providerSettings.sync.editRegionface
    providerSettings.sync.emotionPrompt(default)
    providerSettings.sync.occlusionDetectionEnabledfalse
    safetyfalse
    safety.checkContentfalse
    safety.modefast
    providerSettings.sync.syncModebounce
    providerSettings.sync.temperature0.50
    ttlfalse
    ttl_value60
    outputFormatMP4
    outputQuality95

    react-1, from the company Sync, does something you can't get from an open-weights video model on a consumer GPU: it takes an existing video of a person and re-animates their facial performance to match a different audio track. New dialogue, new emotional delivery, same face, same camera move. That's the whole node. Two required inputs - video and audio, each a UUID or URL - and a VIDEO comes back.

    Let's be clear about what this is: it's performance transfer, the thing people call "dubbing that looks native" and the thing that also makes deepfakes trivially easy. It's a hosted API model, and Runware's content safety is the only thing standing between you and abuse of someone's likeness. Use it on your own footage or with permission; the tooling gives you rope.

    What you set

    • video / audio - required, as UUIDs or URLs (Runware's upload nodes will stage these for you). The video is the appearance and motion source; the audio is the new performance you're driving it with.
    • providerSettings.sync.editRegion - what moves: face, lips, or head. face is the broad default, head adds head motion for bigger deliveries. This is your first quality lever when the output looks "pasted on."
    • providerSettings.sync.emotionPrompt - happy, sad, angry, disgusted, surprised, neutral, or default. You're not just matching the audio; you can push the emotional register it's read with. "Angry" on the same line can sell a scene that flat lip-sync would kill.
    • providerSettings.sync.temperature - expressiveness of the facial movement, 0–1, default 0.5. Low is stiff and clinical; high is theatrical. Where most people get burned: they leave this at the default and blame the model for dead-eyed output when cranking it to 0.7–0.8 fixes the "uncanny" feel.
    • providerSettings.sync.syncMode - what to do when audio and video lengths don't match: bounce (default, ping-pongs), loop, cut_off, silence, or remap. remap is the one that stretches or squeezes the video to fit the audio; the others fill or trim.
    • providerSettings.sync.activeSpeakerDetection.autoDetect - automatically find the talking person in a multi-person shot. Turn it on when there's more than one face.
    • providerSettings.sync.occlusionDetectionEnabled - handles faces that are briefly covered (hand over mouth, passing object). On for talking-head footage where people gesture; it buys robustness at some latency.

    numberResults (up to 4) gives you takes, outputFormat is MP4/WEBM/MOV, and safety.mode (none/fast/full) controls the content gate.

    How it works

    The node sends sync:react-1@1 as a videoInference task through the Runware SDK - your audio and video are referenced by their hosted URLs, the API does the performance transfer, and the finished clip comes back as a native VIDEO socket. It runs entirely server-side; nothing is computed on your machine, and the title bar prints the cost and NSFW verdict when it's done.

    Install and gotchas

    cd ComfyUI/custom_nodes
    git clone https://github.com/Runware/ComfyUI-Runware
    pip install -r ComfyUI-Runware/requirements.txt
    

    Restart; API key from the Runware dashboard into Settings (or RUNWARE_API_KEY). It's paid, per-run, minimum top-up.

    Real-world pain points: it's not a fast node - performance transfer over the API takes a while, so batch sensibly. The temperature default really is the difference between watchable and uncanny, so touch it before you touch anything else. And the input format trips people up: both inputs are strings (UUID or URL), not file paths or IMAGE/AUDIO sockets, so stage media through Runware's upload nodes first and paste the reference in.

    CategoryRunware/Video/sync

    Inputs (16)

    NameTypeDefaultDescription
    videoSTRINGVideo input (UUID or URL).
    audioSTRINGAudio input (UUID or URL).
    numberResultsoptINT11–4Number of results to generate. Each result uses a different seed, producing variations of the same parameters.
    providerSettings.sync.activeSpeakerDetectionoptBOOLEANfalseAutomatically target the speaking face in a multi-person clip.
    providerSettings.sync.editRegionoptCOMBOfaceRegion of the subject to animate during re-animation.
    providerSettings.sync.emotionPromptoptCOMBO(default)Emotional tone for performance re-animation.
    providerSettings.sync.occlusionDetectionEnabledoptBOOLEANfalseEnable occlusion handling for obstructed faces.
    safetyoptBOOLEANfalseEnable to set safety. Off uses the model's default.
    safety.checkContentoptBOOLEANfalseEnable or disable content safety checking. Increases total generation time.
    safety.modeoptCOMBOfastSafety checking mode for video generation.
    providerSettings.sync.syncModeoptCOMBObounceSynchronization strategy when audio and video durations don't match.
    providerSettings.sync.temperatureoptFLOAT0.500–1Expressiveness of lip sync and facial movements.
    ttloptBOOLEANfalseEnable to set ttl. Off uses the model's default.
    ttl_valueoptINT60Time-to-live (TTL) in seconds for generated content. Only applies when `outputType` is `URL`.
    outputFormatoptCOMBOMP4File format for the generated video.
    outputQualityoptINT9520–99Compression quality of the output. Higher values preserve quality but increase file size.

    Outputs (1)

    NameTypeDescription
    videoVIDEO