Open Source RAG Projects
Discover 51 open source RAG repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. RAG projects here are most often combined with Python, LLM and AI. Last updated October 3, 2026.
51 repositories · updated October 3, 2026

PaddleOCR: Extract Text and Structure from Images and PDFs
PaddleOCR is a Python toolkit for recognizing text and parsing document layouts in images and PDFs. It suits developers building OCR, document-processing, and retrieval workflows who need structured output and multilingual recognition.

Intervo: Open-Source Conversational AI Platform for Voice and Chat
Intervo is an open-source platform designed for building, deploying, and managing advanced, goal-oriented AI agents for both voice and chat. It enables users to create complex, multi-step conversational workflows that understand user intent, perform tasks, and integrate seamlessly with existing systems. This versatile platform supports multimodal interactions, from real-time voice calls to web chat, making it suitable for a wide range of applications.

NUDGE: Lightweight Non-Parametric Embedding Fine-Tuning for Retrieval
NUDGE is a lightweight, non-parametric tool designed to fine-tune pre-trained embeddings, significantly enhancing retrieval and RAG pipelines. It operates by adjusting data embeddings directly, rather than modifying model parameters, to maximize accuracy. This approach often leads to over 10% improvement in retrieval accuracy and runs in minutes.

vanna: Turn Natural-Language Questions into SQL Insights
Vanna helps teams build chat interfaces that translate natural-language questions into SQL and return streamed tables, charts, and summaries. It is aimed at applications that need database-backed answers with user-aware permissions and an embedded web UI.

AI Engineering Toolkit: 100+ Libraries for LLM Development
The AI Engineering Toolkit is a comprehensive, curated list featuring over 100 libraries and frameworks essential for AI engineers. It provides battle-tested tools, frameworks, and reference implementations to develop, deploy, and optimize applications built with Large Language Models. This resource aims to help engineers build better LLM apps faster, smarter, and production-ready.

RAG-Anything: Search Documents Across Text, Images, Tables, and Equations
RAG-Anything extends LightRAG with a pipeline for parsing and querying multimodal documents. It is aimed at teams and researchers who need one retrieval system for mixed-content files rather than text-only RAG.

giskard-oss: Test and Red-Team LLM Agents
Giskard is a Python toolkit for evaluating agent behavior and probing AI systems for vulnerabilities. It suits teams building LLM agents or RAG applications that need repeatable checks, safety testing, and adversarial evaluation.

turboseek: An Open-Source AI Search Engine Inspired by Perplexity
turboseek is an innovative open-source AI search engine developed by Nutlope, drawing inspiration from platforms like Perplexity. Built with TypeScript, it leverages advanced LLMs and search APIs to provide comprehensive answers and related follow-up questions. This project offers a robust foundation for anyone interested in building their own AI-powered search solution.

obsidian-smart-composer: Add Vault-Aware AI Writing to Obsidian
Smart Composer is an Obsidian plugin for chatting with AI using selected vault notes, searching notes semantically, and applying suggested edits. It suits writers and researchers who want AI assistance inside their note workflow, with hosted or local model options.

xberg: Extract Text and Structure from Documents
Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

Memary: Add Memory and Knowledge Graphs to AI Agents
Memary is a memory layer for autonomous agents that combines a knowledge graph with user-focused memory to inform responses over time. It suits developers building personalized agents who can manage local models, database connections, and API credentials.

Memori: Give AI Agents Persistent Memory
Memori adds structured, persistent memory to LLM applications by capturing agent execution and conversations. Its Python and TypeScript SDKs integrate with existing models and data infrastructure, with managed cloud and BYODB options.