Open Source Audio Projects
Audio technology covers the creation, playback, recording, processing, and analysis of sound. Tools in this area help people manage music libraries, edit recordings, visualize sound, transcribe speech, and generate music or spoken audio. They can support creative work, accessibility, communication, research, and everyday listening across desktop, web, and mobile environments.
Open source audio tools range from media players and signal-processing libraries to speech recognition, synthesis, and machine-learning applications. When choosing one, consider its license, development activity, documentation, supported formats, hardware and operating system requirements, privacy practices, and compatibility with your existing workflows. These tools are useful for musicians, developers, researchers, educators, and anyone who wants more control over how sound is created or handled.
5 repositories · updated October 3, 2026

VoxCPM: Generate and Clone Multilingual Speech
VoxCPM is a tokenizer-free text-to-speech system for multilingual speech generation, voice design, and voice cloning. Its VoxCPM2 release targets teams and developers who need expressive speech synthesis and can support a 2B-parameter model.

MOSS-TTS: Generate Speech, Dialogue, and Sound with AI
MOSS-TTS is a family of speech and sound generation models for long-form narration, voice cloning, dialogue, voice design, sound effects, and streaming TTS. It offers multiple model architectures and deployment paths for research, production, and local inference.

kapre: Add Audio Preprocessing Layers to Keras Models
Kapre provides TensorFlow-backed Keras layers for transforming audio waveforms into STFT, mel-spectrogram, and related features inside a model. It suits audio ML developers who want preprocessing to travel with training and deployment.

Magenta RT: Live Music Generation on Your Local Device
Magenta RealTime (Magenta RT) is an open-source Python library for live music audio generation on local devices. It allows users to create music using both text and audio prompts, serving as a powerful tool for real-time creative audio exploration. This library is the on-device companion to Google's MusicFX DJ Mode and the Lyria RealTime API.

riffusion-hobby: Generate Music Audio with Stable Diffusion
Riffusion-hobby is a Python library and application for generating music and audio with stable diffusion, including prompt interpolation and spectrogram-to-audio conversion. It suits developers and experimenters who want to run inference locally, but is no longer actively maintained.