Open Source RAG Projects
Discover 54 open source RAG repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. RAG projects here are most often combined with Python, AI Agents and LLM. Last updated October 3, 2026.
54 repositories · updated October 3, 2026

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs
Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

e2m: Convert Documents and Media into Markdown
E2M is a Python library for parsing documents, web pages, and audio into Markdown through configurable parser and converter components. It is aimed at developers preparing varied source material for RAG, training data, or downstream text workflows.

local-deep-research: Research with Local or Cloud AI
Local Deep Research turns complex questions into cited reports by coordinating LLMs with web, academic, and private-document search. It suits researchers and privacy-conscious teams who want a self-hostable tool and control over models and data.

airweave: Retrieve Shared Context for AI Agents
Airweave syncs data from connected apps, databases, and documents into a unified search layer for AI agents and RAG systems. It suits teams that want reusable retrieval infrastructure instead of building separate ingestion pipelines for each agent.

graphrag: Build Knowledge Graphs for LLM Question Answering
Microsoft GraphRAG is a Python pipeline that uses LLMs to turn unstructured text into structured, graph-based context for question answering. It is suited to teams exploring graph-enhanced retrieval over private data, with indexing costs and maintenance-mode status to consider.

opik: Trace, Evaluate, and Monitor LLM Applications
Opik is a platform for tracing and evaluating LLM applications, RAG systems, and AI agents. Teams can use it to inspect workflows, run evaluations, and monitor deployments, either self-hosted or through Comet Cloud.

GenerativeAICourse: A Comprehensive Hands-On Generative AI Engineering Course
This repository offers a comprehensive, hands-on Generative AI course, starting from fundamental AI concepts to building production-grade applications. It focuses on AI engineering, covering topics like LLMs, RAG, AI agents, and prompt engineering with practical tutorials. The course aims to equip learners with the skills needed to build real-world AI solutions.

KAG: Build Knowledge-Grounded Reasoning and Q&A Systems
KAG is a Python framework for building domain-specific question-answering systems that combine knowledge graphs, source text, and LLMs. It targets factual and multi-hop reasoning where vector similarity alone may be insufficient.

attachments: Turn Files Into LLM-Ready Context
attachments is a Python library and CLI that turns documents, images, audio, and other inputs into structured text and image artifacts for LLM workflows. It supports local processing, optional service fallback, and adapters for prompts, chat APIs, and RAG chunks.

kotaemon: Chat with Documents Using RAG
kotaemon is a customizable web interface and toolkit for asking questions about documents with retrieval-augmented generation. It suits teams that want a self-hosted document QA app or developers building and adapting RAG pipelines.

data-prep-kit: Prepare Data for LLM Applications
Data-Prep-Kit is a toolkit for cleaning, transforming, and enriching unstructured data used in LLM training and RAG pipelines. It offers reusable transforms that run with Python or Ray, from local experiments to larger-scale processing.

rag-web-ui: Build Knowledge-Base Q&A with RAG
RAG Web UI is a self-hostable web application for creating question-answering systems over your own documents. It combines document ingestion, vector search, and configurable cloud or local language models, with a web interface and OpenAPI access.