chatterbox-vllm vs csm
Chatterbox vLLM and CSM speech generation compared
chatterbox-vllm adapts Chatterbox for vLLM-based, batched speech-token generation, while csm generates speech from text and optional conversational audio context. The main difference is that chatterbox-vllm focuses on inference through vLLM, whereas csm provides conversational speech generation using a Llama backbone and Mimi audio codes.

chatterbox-vllm: Generate Speech with Chatterbox on vLLM
A vLLM port of the Chatterbox text-to-speech model, built to improve GPU throughput and support batched generation. It suits developers with compatible Nvidia hardware who can work with an early, changing implementation.

csm: Generate Conversational Speech from Text and Audio
CSM is Sesame’s speech-generation model, producing audio from text and optional conversation context. It suits developers building voice experiences who can run large models on a CUDA-compatible GPU and provide the required Hugging Face checkpoints.
| chatterbox-vllm | csm | |
|---|---|---|
| Language | Python | Python |
| License | MIT | Apache-2.0 |
| Stars | 385 | 14.7k |
| Forks | 64 | 1.5k |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- chatterbox-vllm targets batched Chatterbox inference on Nvidia GPUs; csm supports speech generation conditioned on prior utterances, transcripts, and speaker identities.
- chatterbox-vllm uses vLLM for speech-token generation and Chatterbox's S3Gen for waveforms; csm combines a Llama backbone with a Mimi audio decoder.
- chatterbox-vllm is MIT-licensed; csm is Apache-2.0-licensed.
- chatterbox-vllm relies on vLLM internal APIs and is described as likely to work only with vLLM 0.9.2; csm provides code for running its 1B model and requires access to Hugging Face checkpoints.
- chatterbox-vllm has early multilingual support with known quality issues; csm's README cautions that non-English performance is likely limited.
- chatterbox-vllm is aimed at developers testing throughput and batching; csm suits builders adding generated speech to applications that provide the text.
Choose chatterbox-vllm if you…
- need to test batched Chatterbox speech generation through vLLM on supported Nvidia hardware.
- want to compare vLLM inference with the original Chatterbox implementation.
- are exploring audio-conditioned generation or exaggeration controls and can work with an early implementation.
Choose csm if you…
- need speech generation conditioned on conversation context, transcripts, or multiple speaker identities.
- want a Llama-based speech-generation component for application-provided text.
- can run the model with a CUDA-compatible GPU and obtain the required Hugging Face checkpoints.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.