Open Source Text-to-Speech Projects
Discover 31 open source Text To Speech repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Text To Speech projects here are most often combined with Python, AI and Machine Learning. Last updated October 3, 2026.
31 repositories · updated October 3, 2026

BrowserAI: Run AI Models Directly in Your Browser
BrowserAI is a TypeScript library for running language, speech, and audio models locally in a web browser. It suits developers building privacy-conscious AI features without server-side inference, provided users have a compatible WebGPU browser and hardware.

InfiniteTalk: Generate Audio-Driven Talking Videos
InfiniteTalk generates talking videos from audio and an image or existing video, synchronizing speech with facial expressions and body movement. It is aimed at creators and developers who need dubbed or audio-driven video, including long-form generation.

podcastfy: Turn Multimodal Sources Into AI Podcasts
Podcastfy is a Python package and CLI for turning websites, PDFs, images, YouTube videos, and topics into multilingual conversational audio. It suits developers and creators who want to customize or automate podcast generation with hosted or local language models.

captcha: Generate Image and Audio CAPTCHAs
captcha is a Python library for generating image and audio CAPTCHAs, with built-in voice and font data and support for custom assets. It suits applications that need to create their own CAPTCHA challenges and save them as image or audio files.

xiaozhi-esp32-server: Run a Backend for ESP32 Voice Devices
A self-hosted backend for xiaozhi-esp32 devices, providing voice interaction, AI model integrations, device control, and an administration console. It suits ESP32 owners who want to deploy and configure their own service.

ChatTTS: Generate Expressive Speech for Dialogue
ChatTTS is a generative text-to-speech model built for dialogue, including assistant-style speech. It supports Chinese and English, multiple speakers, and controls for prosody such as pauses and laughter.

chatterbox-vllm: Generate Speech with Chatterbox on vLLM
A vLLM port of the Chatterbox text-to-speech model, built to improve GPU throughput and support batched generation. It suits developers with compatible Nvidia hardware who can work with an early, changing implementation.