Speech-to-Text Tools
Speech-to-text technology converts spoken audio into written words, using speech recognition models to identify language and produce transcripts. It can make recordings searchable, support dictation and captions, and enable voice-driven interfaces. Some systems process complete audio files, while others transcribe speech as it arrives, which is useful when low latency matters. Accuracy can vary with language, accents, background noise, recording quality, and specialized vocabulary.
Open source tools in this area include recognition models, inference libraries, transcription applications, and components for live audio pipelines. When choosing one, consider language coverage, accuracy for your audio, offline and hardware requirements, latency, license terms, maintenance activity, and integration with your existing workflow. These tools are useful to developers building voice features, researchers working with audio, and individuals or organizations that need transcription, captions, or local control over speech data.
2 repositories · updated June 25, 2026

Voicebox: The Open-Source AI Voice Studio for Cloning and Dictation
Voicebox is an innovative open-source AI voice studio that allows users to clone voices, generate speech in multiple languages, and dictate into any application. It provides a comprehensive, local-first voice I/O stack, offering a powerful alternative to cloud-based solutions. This tool ensures complete privacy and control over your voice data, running entirely on your local machine.

audapolis: An Editor for Spoken-Word Audio with Automatic Transcription
audapolis is an innovative editor designed for spoken-word audio, offering automatic transcription capabilities. It provides a word processor-like experience for media editing, making the workflow for podcasts, audiobooks, and interviews more efficient. This free and open-source tool keeps your data local, ensuring privacy and control.