Open Source Text-to-Speech Projects

Text-to-speech (TTS) technology converts written text into spoken audio. It can make digital content more accessible, support hands-free use, and provide spoken output for applications such as virtual assistants, language-learning tools, and narration. Modern systems can produce natural-sounding speech, vary tone and pacing, and work across different languages and voices. Some run locally, while others use hosted services or APIs.

Open source options include speech synthesis models, voice-cloning and conversion tools, offline libraries, and interfaces for integrating speech into applications. When choosing one, consider output quality, language and voice support, hardware and runtime requirements, license terms, maintenance activity, and compatibility with your existing workflow. These tools are useful to developers, researchers, accessibility teams, and creators who need to generate or customize spoken audio.

31 repositories · updated October 3, 2026

Speech: Build and Deploy Speech AI Models

Speech: Build and Deploy Speech AI Models

NVIDIA NeMo Speech is a Python framework for researchers and developers building speech recognition, text-to-speech, and speech language models. Use it to train, customize, and run speech models with PyTorch and pretrained checkpoints.

PythonAIMachine Learning
Added Jul 12, 2026 View details
voicebox: Create and Use AI Voices Locally

voicebox: Create and Use AI Voices Locally

Voicebox is a desktop voice studio for local speech generation, voice cloning, dictation, and agent audio. It suits creators and developers who want voice tools and integrations on their own machine, with hardware and platform capabilities varying by setup.

AITypeScriptDesktop App
Added Jun 25, 2026 View details
Open-LLM-VTuber: Talk to LLMs with a Live2D Avatar

Open-LLM-VTuber: Talk to LLMs with a Live2D Avatar

Open-LLM-VTuber is a cross-platform AI companion for voice conversations with language models, paired with a customizable Live2D avatar. It supports local and hosted models, with web and desktop modes for users who want a more expressive, hands-free interface.

PythonAILLM
Added Jun 14, 2026 View details
VoxCPM: Generate and Clone Multilingual Speech

VoxCPM: Generate and Clone Multilingual Speech

VoxCPM is a tokenizer-free text-to-speech system for multilingual speech generation, voice design, and voice cloning. Its VoxCPM2 release targets teams and developers who need expressive speech synthesis and can support a 2B-parameter model.

PythonMachine LearningDeep Learning
Added Jun 1, 2026 View details
MOSS-TTS: Generate Speech, Dialogue, and Sound with AI

MOSS-TTS: Generate Speech, Dialogue, and Sound with AI

MOSS-TTS is a family of speech and sound generation models for long-form narration, voice cloning, dialogue, voice design, sound effects, and streaming TTS. It offers multiple model architectures and deployment paths for research, production, and local inference.

PythonMachine LearningGenerative AI
Added May 31, 2026 View details
tortoise.cpp: Generate Speech with Tortoise TTS in C++

tortoise.cpp: Generate Speech with Tortoise TTS in C++

tortoise.cpp is a C++ implementation of Tortoise TTS built on ggml for local text-to-speech generation. It supports CPU and CUDA builds, with Metal support described as in progress.

C++Text To SpeechMachine Learning
Added May 12, 2026 View details
StreamingKokoroJS: Generate Speech Locally in Your Browser

StreamingKokoroJS: Generate Speech Locally in Your Browser

StreamingKokoroJS turns text into speech in the browser with the Kokoro-82M model. It streams audio locally, with WebGPU acceleration when available and a WebAssembly fallback, without sending text to a server.

JavaScriptAIText To Speech
Added Apr 26, 2026 View details
chatterbox: Generate Speech from Text with Voice Cloning

chatterbox: Generate Speech from Text with Voice Cloning

Chatterbox is Resemble AI’s Python text-to-speech model family, with English and multilingual options, voice cloning, and expressive speech controls. It suits developers building voice experiences who can manage model inference and its compute requirements.

PythonAIMachine Learning
Added Apr 19, 2026 View details
pocketpal-ai: Run Language Models on Your Phone

pocketpal-ai: Run Language Models on Your Phone

PocketPal AI is a mobile assistant that runs language models and text-to-speech on-device, so chats can stay private and work offline. It suits people who want local AI on iOS or Android without relying on a cloud service.

AILLMTypeScript
Added Apr 15, 2026 View details
Spark-TTS: Generate Speech and Clone Voices from Text

Spark-TTS: Generate Speech and Clone Voices from Text

Spark-TTS is a PyTorch inference project for bilingual text-to-speech and zero-shot voice cloning. It uses a Qwen2.5-based model to generate speech and supports adjustable voice characteristics.

PythonText To SpeechMachine Learning
Added Apr 5, 2026 View details
RedditVideoMakerBot: Generate Videos from Reddit Posts

RedditVideoMakerBot: Generate Videos from Reddit Posts

RedditVideoMakerBot automates the creation of short videos from Reddit content, without manual video editing or asset compiling. It is for creators who want to generate a video file for manual upload to social platforms.

PythonAutomationCLI
Added Mar 30, 2026 View details
index-tts-lora: Fine-Tune IndexTTS for Custom Voices

index-tts-lora: Fine-Tune IndexTTS for Custom Voices

A Python project that adds LoRA fine-tuning workflows to IndexTTS for single- and multi-speaker voice synthesis. It is aimed at users who want to adapt speech generation to speaker audio and improve prosody and naturalness.

PythonAIMachine Learning
Added Mar 23, 2026 View details
Previous Page 1 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️