
InfiniteTalk: Audio-Driven Full-Body Video Dubbing turns a single image and one or more audio tracks into a lip‑synced, full‑body video while preserving the original identity, background, and camera motion. The core of the pipeline is WanInfiniteTalkToVideo, which fuses features from the loaded Wan2.1 I2V model with an audio embedding to drive body and facial motion. Step1 - Load Models prepares the UNet and VAE (UNETLoader, VAELoader) for image‑to‑video synthesis, and loads text and audio backbones (CLIPLoader with CLIPTextEncode for camera/stage prompts, and AudioEncoderLoader with AudioEncoderEncode for voice features). You import the starting frame via LoadImage, then draw masks per character in the MaskEditor to localize motion to each speaker.
API
Use this workflow from code
Every Comfy workflow is a JSON graph. The payload below is this workflow, exactly as ComfyUI runs it — fetch it from the URL, keep it in version control, or load it in ComfyUI and run it node by node.
GET https://comfy.org/workflows/download/b95499a55c9f.jsonFetching workflow JSON…
Run it from Python or TypeScript with the Comfy SDK. The same code targets Comfy Cloud or a ComfyUI you host yourself — only the base URL changes.
# Install (beta)
pip install comfy-sdk # Python
npm i @comfyorg/sdk # TypeScript
# Run "InfiniteTalk: Audio-Driven Full-Body Video Dubbing" (Python)
from comfy_sdk import Comfy
client = Comfy(api_key="comfyui-...")
# This workflow, exported in API format (see note below)
wf = client.workflows.from_file("video_wan2_1_infinitetalk_api.json")
asset = client.assets.from_file("input.png")
wf.set_input("32", "image", asset) # LoadImage
job = client.run(wf)
for output in job.get_outputs("162"): # SaveVideo
output.to_file(output.name)The SDK takes a workflow in API format: open this workflow in ComfyUI and use File → Export Workflow (API). API access requires a Comfy Cloud plan with an API key. SDK docs · Get an API key
FAQ














