Open Source Text-to-Speech Projects
Text-to-speech (TTS) technology converts written text into spoken audio. It can make digital content more accessible, support hands-free use, and provide spoken output for applications such as virtual assistants, language-learning tools, and narration. Modern systems can produce natural-sounding speech, vary tone and pacing, and work across different languages and voices. Some run locally, while others use hosted services or APIs.
Open source options include speech synthesis models, voice-cloning and conversion tools, offline libraries, and interfaces for integrating speech into applications. When choosing one, consider output quality, language and voice support, hardware and runtime requirements, license terms, maintenance activity, and compatibility with your existing workflow. These tools are useful to developers, researchers, accessibility teams, and creators who need to generate or customize spoken audio.
31 repositories · updated October 3, 2026

Speech: Build and Deploy Speech AI Models
NVIDIA NeMo Speech is a Python framework for researchers and developers building speech recognition, text-to-speech, and speech language models. Use it to train, customize, and run speech models with PyTorch and pretrained checkpoints.

voicebox: Create and Use AI Voices Locally
Voicebox is a desktop voice studio for local speech generation, voice cloning, dictation, and agent audio. It suits creators and developers who want voice tools and integrations on their own machine, with hardware and platform capabilities varying by setup.

Open-LLM-VTuber: Talk to LLMs with a Live2D Avatar
Open-LLM-VTuber is a cross-platform AI companion for voice conversations with language models, paired with a customizable Live2D avatar. It supports local and hosted models, with web and desktop modes for users who want a more expressive, hands-free interface.

VoxCPM: Generate and Clone Multilingual Speech
VoxCPM is a tokenizer-free text-to-speech system for multilingual speech generation, voice design, and voice cloning. Its VoxCPM2 release targets teams and developers who need expressive speech synthesis and can support a 2B-parameter model.

MOSS-TTS: Generate Speech, Dialogue, and Sound with AI
MOSS-TTS is a family of speech and sound generation models for long-form narration, voice cloning, dialogue, voice design, sound effects, and streaming TTS. It offers multiple model architectures and deployment paths for research, production, and local inference.

tortoise.cpp: Generate Speech with Tortoise TTS in C++
tortoise.cpp is a C++ implementation of Tortoise TTS built on ggml for local text-to-speech generation. It supports CPU and CUDA builds, with Metal support described as in progress.

StreamingKokoroJS: Generate Speech Locally in Your Browser
StreamingKokoroJS turns text into speech in the browser with the Kokoro-82M model. It streams audio locally, with WebGPU acceleration when available and a WebAssembly fallback, without sending text to a server.

chatterbox: Generate Speech from Text with Voice Cloning
Chatterbox is Resemble AI’s Python text-to-speech model family, with English and multilingual options, voice cloning, and expressive speech controls. It suits developers building voice experiences who can manage model inference and its compute requirements.

pocketpal-ai: Run Language Models on Your Phone
PocketPal AI is a mobile assistant that runs language models and text-to-speech on-device, so chats can stay private and work offline. It suits people who want local AI on iOS or Android without relying on a cloud service.

Spark-TTS: Generate Speech and Clone Voices from Text
Spark-TTS is a PyTorch inference project for bilingual text-to-speech and zero-shot voice cloning. It uses a Qwen2.5-based model to generate speech and supports adjustable voice characteristics.

RedditVideoMakerBot: Generate Videos from Reddit Posts
RedditVideoMakerBot automates the creation of short videos from Reddit content, without manual video editing or asset compiling. It is for creators who want to generate a video file for manual upload to social platforms.

index-tts-lora: Fine-Tune IndexTTS for Custom Voices
A Python project that adds LoRA fine-tuning workflows to IndexTTS for single- and multi-speaker voice synthesis. It is aimed at users who want to adapt speech generation to speaker audio and improve prosody and naturalness.