Open Source Text-to-Speech (TTS) Tools
Text-to-speech (TTS) technology converts written text into spoken audio. It supports accessibility, spoken interfaces, narration, language learning, and automated voice responses, reducing the effort needed to produce speech at scale. Modern systems can vary pronunciation, pacing, and expression, while some can adapt speech to a particular voice or speaking style. They may run locally or provide speech generation through a service or application programming interface.
Open source TTS tools include pretrained speech models, training and fine-tuning frameworks, voice conversion systems, and libraries for integrating synthesis into applications. When choosing a tool, consider voice quality and language coverage, licensing and voice-use rights, maintenance activity, hardware and runtime requirements, and compatibility with existing workflows. These tools are useful to developers, researchers, accessibility teams, and creators who need control over how speech is generated and deployed.
3 repositories · updated July 12, 2026

NVIDIA NeMo Speech: Scalable Generative AI for Speech Models
NVIDIA NeMo Speech is a powerful, scalable generative AI framework designed for researchers and developers focused on Large Language Models, Multimodal, and Speech AI. It provides tools for Automatic Speech Recognition (ASR) and Text-to-Speech (TTS), enabling efficient creation, customization, and deployment of new AI models using existing code and pre-trained checkpoints. This framework supports a wide range of applications, from real-time streaming ASR to high-quality multilingual TTS.

StreamingKokoroJS: Unlimited, Local Text-to-Speech in Your Browser
StreamingKokoroJS provides unlimited text-to-speech capabilities directly within your browser, ensuring 100% local processing and complete privacy. This open-source project leverages the Kokoro-JS model and WebGPU acceleration to deliver high-quality, streaming audio generation without server-side interaction.

chatterbox-vllm: Accelerating Chatterbox TTS with vLLM for Enhanced Performance
chatterbox-vllm is a high-performance port of the Chatterbox Text-to-Speech (TTS) model to vLLM, designed to significantly improve generation speed and GPU memory efficiency. This personal project aims to provide a more efficient and easily integratable solution for speech synthesis, offering substantial speedups compared to the original implementation. While currently usable and demonstrating benchmark-topping throughput, it leverages internal vLLM APIs and hacky workarounds, with ongoing refactoring planned.