ComfyUI Node

MuseVRun

Animate one photo into a natural virtual-human clip with Tencent's MuseV

By chaojie·Created 2 years ago·Updated 2 years ago· 156
MuseVRun
  • ref_image
  • IMAGE
prompt(masterpiece, best quality, highres:1),(1girl, solo:1),(beautiful face, soft skin, costume:1),(eye blinks:1.8),(head wave:1.3)
img_edge_ratio1.00
video_len12
seed1234
width563
height714
first_redraw_steps30
video_steps10
first_redraw_guidance_scale7.50
video_guidance_scale3.50
sd_model_namemajicmixRealv6Fp16
uselcmfalse

MuseVRun is the ComfyUI wrapper around Tencent's MuseV, the model behind half the "AI influencer" clips you saw in spring 2024. You feed it one portrait and it generates a short clip of that person moving - subtle head motion, eye blinks, a slight turn. If you want a talking-head avatar or a virtual-human asset, this is the node. If you want general video generation, don't come here: Wan and Hunyuan long ago outclassed it on arbitrary scenes. MuseV's niche is specifically virtual humans on a small SD 1.5 footprint, and in that niche it still has no real rival.

What it is

MuseV is a human-specific image-to-video model released by Tencent's TMElyralab team in March 2024. It rides the SD 1.5 ecosystem: a motion module plus a reference network (musev_referencenet) and IP-Adapter give the frozen base model temporal awareness and appearance lock-in. The pack's author, chaojie, did the minimal thing - the README is a paragraph and two workflow files, and the real documentation lives in the upstream TMElyralab/MuseV repo. Don't let the thin README worry you; the node works.

The output is intentionally small-motion: eye blinks, head waves, breathing. The default prompt ((eye blinks:1.8), (head wave:1.3)) tells you exactly what the authors trained it for. Fixed-camera character shots, not camera moves.

How it works

The node saves your ref_image to a temp file, then runs a two-phase pipeline. First a first-frame redraw: an img2img pass (first_redraw_steps, default 30) cleans up your input and establishes the look. Then the actual video pass runs parallel denoising - the motion module generates all frames together rather than frame-by-frame, which is why it avoids the error accumulation that plagued early AnimateDiff workflows. Appearance comes from the reference net + IP-Adapter keyed on your image, motion comes from the prompt and the training.

One honest gotcha from the source: video_len maps to time_size, the number of frames generated per batch (default 12, and the shipped workflow cranks it to 120). Each generated frame is also capped in motion - longer and wider means smaller motion, as the config comments note.

Inputs that matter

  • ref_image (IMAGE) - your portrait. Feed it straight from LoadImage.
  • prompt (STRING) - drives motion and style. The default is a good starting point for human subjects.
  • video_len (INT) - total frames of the clip. 12 is short; 60–120 is typical.
  • first_redraw_steps / first_redraw_guidance_scale - how much the first frame is redrawn (30 / 7.5).
  • video_steps / video_guidance_scale - the video pass (10 / 3.5). Higher guidance = more motion, but more chance of the clip falling apart.
  • sd_model_name - majicmixRealv6Fp16 (default) or fantasticmix_v10, both SD 1.5 checkpoints that ship inside the MuseV model download.
  • uselcm (BOOLEAN) - flips on an LCM-LoRA for much faster generation at lower quality. Good for previews.

The output is a single IMAGE tensor - a batch of frames (T, H, W, C) - which you wire into VHS_VideoCombine to save as an mp4. That's exactly how the shipped workflow is wired. Skip the img_edge_ratio slider; the source hardcodes it to 1.0 in this version, so it's inert.

Install

Grab it from ComfyUI Manager (search "ComfyUI-MuseV"), or manually:

cd ComfyUI/custom_nodes
git clone --recursive https://github.com/chaojie/ComfyUI-MuseV
cd ComfyUI-MuseV
pip install -r requirements.txt

The --recursive matters: this pack vendors three submodules (MMCM, controlnet_aux, a diffusers fork), and a plain clone gets you an empty MMCM/. Then the big one - the model, several gigabytes:

huggingface-cli download --resume-download TMElyralab/MuseV --local-dir ComfyUI/models/diffusers/TMElyralab/MuseV

That one download contains everything: the motion UNet, the SD 1.5 base checkpoints, IP-Adapter weights, the LCM LoRA, and the negative embeddings. The pack's requirements.txt is short (decord, moviepy, h5py, …), but in practice you also need the pose/OpenMMLab stack (openmim, mmengine, mmcv==2.0.1, mmdet==3.1.0, mmpose==1.1.0) or the node refuses to import. Budget for a beefy GPU too - the community reported MuseV eating a full 24 GB back in 2024.

Common issues

  • "Cannot be loaded" / import errors - almost always the missing MM packages or a clone without submodules. See install above.
  • Model not found - the diffusers checkpoint must be at exactly ComfyUI/models/diffusers/TMElyralab/MuseV.
  • OOM - drop width/height or video_len. 563×714 at 12 frames is a gentle start.
  • Slow - enable uselcm and drop video_steps to 4–6 for iteration, then turn it off for final output.

It's a 2024-era model and it shows - but for one-photo avatar animation it remains the fastest path to something that looks alive.

CategoryMuseV

Inputs (13)

NameTypeDefaultDescription
ref_imageIMAGE
promptSTRING(masterpiece, best quality, highres:1),(1girl, solo:1),(beautiful face, soft skin, costume:1),(eye blinks:1.8),(head wave:1.3)
img_edge_ratioFLOAT1.00
video_lenINT12
seedINT1234
widthINT563
heightINT714
first_redraw_stepsINT30
video_stepsINT10
first_redraw_guidance_scaleFLOAT7.50
video_guidance_scaleFLOAT3.50
sd_model_nameCOMBOmajicmixRealv6Fp162 options: majicmixRealv6Fp16, fantasticmix_v10
uselcmBOOLEANfalse

Outputs (1)

NameTypeDescription
IMAGEIMAGE