MuseTalk: Generate Audio-Synced Talking-Head Videos

Summary
MuseTalk creates lip-synced video from a source video or image and an audio clip using latent-space inpainting. It supports training and inference workflows, with real-time performance reported on a Tesla V100.
At a glance
- Language
- Python
- License
- NOASSERTION
- Stars
- 6.7k
- Forks
- 968
- Added to OSRepos
- December 9, 2025
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
MuseTalk is an audio-driven lip-syncing system for modifying a face in an existing video or image so its mouth movements match supplied audio. It targets creators and developers building dubbed videos or virtual-human pipelines, including workflows that pair it with MuseV.
The model encodes image and audio features and performs a single-step latent-space inpainting operation. This is not a diffusion model, despite its use of a UNet architecture related to Stable Diffusion. The repository includes inference, real-time inference, and training code, as well as pretrained weights.
Key Features
- Generate lip-synced video from a source video, image, or image sequence and an audio file.
- Supports audio in languages including Chinese, English, and Japanese.
- Provides standard and real-time inference scripts for MuseTalk 1.0 and 1.5.
- Includes a Gradio interface for adjusting parameters and previewing a first frame.
- Provides dataset preprocessing and staged training scripts.
- Offers a face-region positioning control that can affect mouth openness.
- Integrates with a virtual-human workflow using MuseV.
Use Cases
- Video creators can dub an existing talking-head clip into another language or voice when matching mouth movement is important.
- Virtual-human developers can add speech animation to generated character videos or still images.
- Researchers can experiment with audio-conditioned face generation and train models using the supplied preprocessing and training workflow.
- Teams building a video production tool can use the inference scripts or Gradio demo to evaluate lip-sync results before integrating the model.
Project Facts
- Language: Python
- License: NOASSERTION
- Stars: 6.7k
- Forks: 968
- Topics: lip-sync, virtualhumans
- Archived: false
Getting Started
The README recommends Python 3.10 and CUDA 11.7, and documents PyTorch, MMLab packages, FFmpeg, and model-weight setup. After installation and weight download, Linux users can run MuseTalk 1.5 inference with:
sh inference.sh v1.5 normal
See the README for Windows instructions, real-time inference, model downloads, and training setup. Model weights are also available from the Hugging Face repository, and a hosted demo is provided.
Alternatives
- audio2photoreal: audio2photoreal generates a photorealistic avatar driven by audio, while MuseTalk lip-syncs an existing video or image to an audio clip.
Considerations
- The documented setup requires a CUDA-capable GPU, several model dependencies, FFmpeg, and separate pretrained weights. The project reports 30fps+ on an NVIDIA Tesla V100, but speed depends on hardware and configuration.
- The README notes limitations in identity preservation and temporal stability, including possible changes to facial details and jitter from single-frame generation.
- The generated face region is 256 × 256. Higher-resolution output may require an additional super-resolution model.
- The repository says its code is MIT-licensed, while the supplied repository metadata reports the license as NOASSERTION. Check the repository license and the terms of third-party models and datasets before use. The README restricts test data to non-commercial research purposes.
Source repository
Open the original repository on GitHub.
27 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.

agent-anvil: Test AI Agent Tool Use in CI
October 1, 2026
Agent Anvil evaluates tool-using AI agents through scenario-based runs, trace checks, and optional semantic grading. It helps teams catch unsafe or incorrect tool behavior before it reaches production.