Open Source Text-to-Speech Projects
Discover 31 open source Text To Speech repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Text To Speech projects here are most often combined with Python, AI and Machine Learning. Last updated October 3, 2026.
31 repositories · updated October 3, 2026

sherpa-onnx: Run Speech AI Locally Across Platforms
sherpa-onnx runs speech and audio models locally through ONNX Runtime, without an internet connection. It supports tasks from speech recognition and synthesis to diarization, with APIs and deployment options spanning mobile, desktop, web, and embedded devices.

Intervo: Build Voice and Chat AI Agents
Intervo is a self-hostable platform for designing conversational AI agents that handle voice calls and web chat. It combines visual workflows, knowledge retrieval, telephony, and integrations with multiple speech and language-model providers.

ttsfm: Provide an OpenAI-Compatible Text-to-Speech API
TTSFM provides a self-hosted text-to-speech service and Python client compatible with OpenAI's speech API. It suits developers who want to generate speech through a REST endpoint, Python, or a web playground, with a disclaimer against production and commercial use.

fastrtc: Build Real-Time Audio and Video Apps in Python
FastRTC turns Python functions into real-time audio and video streams over WebRTC or WebSockets. It suits developers building voice, video, and AI interaction features who want stream handling and optional Gradio or FastAPI integration.

joinly: Let AI Agents Join and Participate in Meetings
Joinly connects AI agents to live video meetings through an MCP server, enabling them to listen, speak, chat, and use tools. It suits developers building meeting-aware agents who want a self-hosted setup and can manage its provider credentials and local runtime.

csm: Generate Conversational Speech from Text and Audio
CSM is Sesame’s speech-generation model, producing audio from text and optional conversation context. It suits developers building voice experiences who can run large models on a CUDA-compatible GPU and provide the required Hugging Face checkpoints.

paper2gui: Run AI Models Through Desktop Apps
Paper2GUI is a desktop toolbox that packages AI models into ready-to-use apps for tasks such as image enhancement, speech synthesis, and video processing. It suits people who want to use these tools without building a development environment.

ElatoAI: Build Realtime Voice AI Devices with ESP32
ElatoAI connects ESP32 devices to realtime voice AI through secure WebSockets and edge functions. It is for developers building talking toys, companions, or other connected devices that need speech-to-speech models and a web-based control app.

youtube-summarizer: Summarize YouTube Videos and Playlists
A Flask web app that turns YouTube video and playlist transcripts into AI-generated summaries, with optional audio playback. It suits individuals or small groups who want a self-hosted way to review video content.

ClickUi: A Desktop Assistant for Chatting with AI Models
ClickUi is a Python desktop assistant that opens with a hotkey and supports text and voice conversations with local or API-based AI models. It suits people who want AI tools available alongside their everyday computer work.

open-notebooklm: Turn PDFs Into Podcast Audio
Open NotebookLM turns a PDF into an AI-generated podcast dialogue and MP3. It suits readers who want an audio-style overview of a document, and requires a Fireworks API key to run.

GPT-SoVITS: Clone Voices and Generate Speech from Text
GPT-SoVITS is a Python toolkit for voice cloning, speech conversion, and text-to-speech. It supports zero-shot synthesis from a short reference clip and fine-tuning with about one minute of voice data, with a WebUI for preparing data and training models.