chatterbox-vllm: Generate Speech with Chatterbox on vLLM

chatterbox-vllm: Generate Speech with Chatterbox on vLLM

Summary

A vLLM port of the Chatterbox text-to-speech model, built to improve GPU throughput and support batched generation. It suits developers with compatible Nvidia hardware who can work with an early, changing implementation.

At a glance

Language
Python
License
MIT
Stars
385
Forks
64
Added to OSRepos
October 11, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

chatterbox-vllm adapts the Chatterbox TTS model to run its speech-token generation through vLLM. The goal is to reduce inference overhead and make batching more efficient than the original Transformers-based implementation.

It is aimed at developers experimenting with Chatterbox inference, especially those generating multiple speech requests on Nvidia GPUs. The project reports promising throughput results, but its vLLM integration relies on internal APIs and is not yet a stable, production-oriented interface.

Key Features

  • Generates speech from text with optional audio prompts for voice conditioning.
  • Supports batched generation through vLLM.
  • Implements Context Free Guidance, configured with the CHATTERBOX_CFG_SCALE environment variable.
  • Provides exaggeration control and an example script for generating audio.
  • Includes early multilingual support, with known quality limitations.
  • Uses Chatterbox's S3Gen implementation for waveform generation alongside vLLM-based speech-token generation.

Use Cases

  • TTS developers comparing vLLM-based inference with the original Chatterbox implementation.
  • Teams generating batches of speech samples on Nvidia GPUs and investigating throughput improvements.
  • Researchers testing audio-conditioned speech cloning and exaggeration controls.
  • Developers exploring multilingual TTS who can tolerate incomplete features and quality issues.

Project Facts

  • Language: Python
  • License: MIT
  • Stars: 385
  • Forks: 64
  • Topics: none listed
  • Archived: No

Getting Started

On Linux or WSL2 with Nvidia hardware, install uv and run:

git clone https://github.com/randombk/chatterbox-vllm.git
cd chatterbox-vllm
uv venv
source .venv/bin/activate
uv sync

See the README for usage examples, benchmark details, and updates.

Alternatives

  • chatterbox: The original Chatterbox implementation offers multilingual synthesis and voice cloning, while chatterbox-vllm focuses on vLLM-based throughput and batching.
  • ChatTTS: ChatTTS focuses on Chinese and English dialogue speech with speaker and prosody controls, rather than vLLM serving for Chatterbox.
  • GPT-SoVITS: GPT-SoVITS combines voice cloning, speech conversion, and TTS with a WebUI and fine-tuning workflows, unlike this Chatterbox vLLM port.
  • csm: CSM generates speech from text with optional conversation context and requires its CUDA-compatible model checkpoints, rather than serving Chatterbox through vLLM.

Considerations

  • Linux and WSL2 with Nvidia hardware are the supported environments. AMD support is untested.
  • The implementation uses vLLM internal APIs and workarounds. The README describes it as likely to work only with vLLM 0.9.2 until integration changes are made.
  • APIs may change, and the server API is not implemented.
  • Learned speech positional embeddings and the Alignment Stream Analyzer are missing. Multilingual output can have errors, repetitions, noise, or other quality degradation.
  • Benchmarks and further performance optimization are incomplete. Benchmark results are hardware- and setup-specific, and the S3Gen waveform stage remains a significant part of generation time.

Comparisons

Source repository

Open the original repository on GitHub.

17 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️