MuseTalk: Generate Audio-Synced Talking-Head Videos

MuseTalk: Generate Audio-Synced Talking-Head Videos

Summary

MuseTalk creates lip-synced video from a source video or image and an audio clip using latent-space inpainting. It supports training and inference workflows, with real-time performance reported on a Tesla V100.

At a glance

Language
Python
License
NOASSERTION
Stars
6.7k
Forks
968
Added to OSRepos
December 9, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

MuseTalk is an audio-driven lip-syncing system for modifying a face in an existing video or image so its mouth movements match supplied audio. It targets creators and developers building dubbed videos or virtual-human pipelines, including workflows that pair it with MuseV.

The model encodes image and audio features and performs a single-step latent-space inpainting operation. This is not a diffusion model, despite its use of a UNet architecture related to Stable Diffusion. The repository includes inference, real-time inference, and training code, as well as pretrained weights.

Key Features

  • Generate lip-synced video from a source video, image, or image sequence and an audio file.
  • Supports audio in languages including Chinese, English, and Japanese.
  • Provides standard and real-time inference scripts for MuseTalk 1.0 and 1.5.
  • Includes a Gradio interface for adjusting parameters and previewing a first frame.
  • Provides dataset preprocessing and staged training scripts.
  • Offers a face-region positioning control that can affect mouth openness.
  • Integrates with a virtual-human workflow using MuseV.

Use Cases

  • Video creators can dub an existing talking-head clip into another language or voice when matching mouth movement is important.
  • Virtual-human developers can add speech animation to generated character videos or still images.
  • Researchers can experiment with audio-conditioned face generation and train models using the supplied preprocessing and training workflow.
  • Teams building a video production tool can use the inference scripts or Gradio demo to evaluate lip-sync results before integrating the model.

Project Facts

  • Language: Python
  • License: NOASSERTION
  • Stars: 6.7k
  • Forks: 968
  • Topics: lip-sync, virtualhumans
  • Archived: false

Getting Started

The README recommends Python 3.10 and CUDA 11.7, and documents PyTorch, MMLab packages, FFmpeg, and model-weight setup. After installation and weight download, Linux users can run MuseTalk 1.5 inference with:

sh inference.sh v1.5 normal

See the README for Windows instructions, real-time inference, model downloads, and training setup. Model weights are also available from the Hugging Face repository, and a hosted demo is provided.

Alternatives

  • audio2photoreal: audio2photoreal generates a photorealistic avatar driven by audio, while MuseTalk lip-syncs an existing video or image to an audio clip.

Considerations

  • The documented setup requires a CUDA-capable GPU, several model dependencies, FFmpeg, and separate pretrained weights. The project reports 30fps+ on an NVIDIA Tesla V100, but speed depends on hardware and configuration.
  • The README notes limitations in identity preservation and temporal stability, including possible changes to facial details and jitter from single-frame generation.
  • The generated face region is 256 × 256. Higher-resolution output may require an additional super-resolution model.
  • The repository says its code is MIT-licensed, while the supplied repository metadata reports the license as NOASSERTION. Check the repository license and the terms of third-party models and datasets before use. The README restricts test data to non-commercial research purposes.

Source repository

Open the original repository on GitHub.

27 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️