Open Source Conversational AI Projects
Conversational AI covers systems that communicate with people through natural language, by text, voice, or a combination of both. These systems can answer questions, guide users through tasks, handle routine support requests, and connect conversations to other services. They bring language processing and dialogue management together to interpret input, maintain context, and produce useful responses, while helping reduce repetitive work and make digital services easier to access.
Open source tools in this area include chatbot frameworks, agent builders, speech recognition and synthesis components, and platforms for real-time voice or video interactions. When choosing one, consider its maturity, license, maintenance activity, hardware and hosting requirements, supported languages, and integration options. Also check how it handles privacy, evaluation, and conversation history. These tools can be useful to developers, researchers, and organizations building support experiences, assistants, or interactive applications.
5 repositories · updated March 19, 2026

Hexabot: Open-Source AI Chatbot and Agent Builder
Hexabot is an open-source AI chatbot and agent builder designed for creating and managing multi-channel and multilingual conversational agents with ease. It offers extensive customization, powerful text-to-action capabilities, and supports integration with various LLM models, making it a flexible solution for developers. This project simplifies the deployment and management of sophisticated AI-powered interactions across different platforms.

Intervo: Open-Source Conversational AI Platform for Voice and Chat
Intervo is an open-source platform designed for building, deploying, and managing advanced, goal-oriented AI agents for both voice and chat. It enables users to create complex, multi-step conversational workflows that understand user intent, perform tasks, and integrate seamlessly with existing systems. This versatile platform supports multimodal interactions, from real-time voice calls to web chat, making it suitable for a wide range of applications.

TEN VAD: Low-Latency, High-Performance Voice Activity Detector
TEN VAD is a low-latency, high-performance, and lightweight Voice Activity Detector (VAD) designed for real-time enterprise use. It provides accurate frame-level speech activity detection, outperforming common alternatives like WebRTC VAD and Silero VAD. This system is crucial for enhancing conversational AI by reducing end-to-end latency and improving speech segment extraction.

joinly: Make Your Meetings Accessible to AI Agents
joinly.ai is an open-source connector middleware designed to integrate AI agents into video calls. It enables agents to actively participate, interact in real-time, and perform tasks during meetings across platforms like Google Meet, Zoom, and Microsoft Teams. The project emphasizes a privacy-first approach, offering self-hosting capabilities and flexibility with various LLM, TTS, and STT providers.

csm: Generate Conversational Speech from Text and Audio
CSM is Sesame’s speech-generation model, producing audio from text and optional conversation context. It suits developers building voice experiences who can run large models on a CUDA-compatible GPU and provide the required Hugging Face checkpoints.