Open Source RAG Projects
Retrieval-augmented generation (RAG) combines information retrieval with language models. A system searches a collection of trusted sources for relevant material, then provides that context to a model as it generates a response. This helps ground answers in specific documents, support knowledge-intensive question answering, and keep information current without relying only on a model’s training data. RAG can also make it easier to trace responses back to source material, though retrieval quality and source coverage affect results.
Open source tools in this area include document parsers, indexing and embedding pipelines, search backends, orchestration frameworks, evaluation tools, and interfaces for querying knowledge bases. When choosing one, consider its license, maintenance activity, supported data formats and models, deployment requirements, integrations, and how well it handles permissions and citations. These tools are useful to developers, researchers, and organizations building searchable assistants or document-based applications.
50 repositories · updated October 3, 2026

Utopia: The Open-Source Enterprise World Model for Knowledge Engineering
Utopia is the world's first open-source enterprise world model, built with Rust. It provides a unique foundation for knowledge engineering, featuring a bitemporal knowledge graph that evolves with new information and supports robust reasoning and decision-making. Designed for self-hosting, Utopia offers companies full control over their knowledge foundation and compliance audit trails.

LLM Wiki: Build a Self-Maintaining, Interlinked Knowledge Base with AI
LLM Wiki is a powerful cross-platform desktop application designed to transform your documents into an organized, interlinked knowledge base automatically. Unlike traditional RAG systems, it incrementally builds and maintains a persistent wiki from your sources, ensuring knowledge is compiled once and kept current. This innovative approach offers a dynamic and evolving personal knowledge management solution.

Guaardvark: Your Self-Hosted AI Studio for Agents, Media, and Code
Guaardvark is a comprehensive, self-hosted AI studio designed for local execution of advanced AI tasks. It integrates coding agents, media generation (video, image, music, voice), and robust RAG capabilities, all running on a single GPU. This platform prioritizes privacy and user control, enabling a full AI workstation experience on your own hardware.

GitNexus: Zero-Server Code Intelligence Engine for AI Agents
GitNexus is a client-side knowledge graph creator that runs entirely in your browser, transforming GitHub repositories or ZIP files into interactive knowledge graphs. It features a built-in Graph RAG Agent, perfect for deep code exploration and enhancing AI agent reliability. This tool provides architectural context to AI, ensuring more accurate and efficient code analysis.

AgentsKit: The Complete JavaScript Toolkit for Building AI Agents
AgentsKit is a comprehensive JavaScript toolkit designed for building AI agents, offering a lightweight core and a modular ecosystem. It provides essential components like UIs, autonomous runtime, tools, memory, and RAG, enabling developers to create sophisticated agents from simple chat interfaces to complex autonomous systems. This framework aims to simplify agent development by offering composable parts and avoiding the need to glue multiple incompatible libraries together.

tooltrim: Drastically Reduce LLM Agent Tool Output Tokens, Improve Accuracy
tooltrim provides drop-in compression for LLM agent tool outputs, drastically cutting tokens while often improving answer accuracy. This provider-agnostic solution offers content-aware compression, faithfulness benchmarks, and seamless integration with popular frameworks or as an OpenAI-compatible proxy.

OpenViking: A Self-Evolving Context Database for AI Agents
OpenViking is an open-source context database designed for AI agents, unifying agent memory, knowledge RAG, and skills into a virtual filesystem. It allows agents to browse their context deterministically using familiar commands like `ls` and `tree`. This innovative approach aims to enhance agent performance and reduce token spend by loading content in tiered layers.

book-to-skill: Transform Technical Books into AI Agent Skills
book-to-skill is a Python project that converts technical books, documents, or source collections into structured agent skills. It allows AI agents like GitHub Copilot CLI or Claude Code to load content on demand, providing accurate answers without hallucination. This tool optimizes learning and reference by distilling complex information into an easily queryable format.

ReMe: Manage Persistent Memory for AI Agents
ReMe turns agent conversations and resources into searchable, editable Markdown memory that can evolve over time. It suits developers building assistants that need reusable knowledge across sessions and agent runtimes.

DeepTutor: Lifelong Personalized Tutoring with AI Agents
DeepTutor is an advanced AI-powered platform designed for lifelong personalized tutoring, integrating various learning modes into a single, extensible system. It leverages large language models and multi-agent systems to offer features like interactive chat, quiz generation, and skill development. This project provides a comprehensive environment for learners and educators seeking intelligent, adaptive educational tools.

Memori: Agent-Native Memory Infrastructure for LLM Production Systems
Memori provides agent-native memory infrastructure, offering an LLM-agnostic layer that transforms agent execution and conversations into structured, persistent state. Designed for enterprise use, it seamlessly integrates with existing data infrastructure and supports various deployment environments, ensuring robust memory management for AI agents.

AI-Agents-Projects-Tutorials: Comprehensive Guide to AI Agent Development
The AI-Agents-Projects-Tutorials repository offers an extensive collection of code implementations and tutorials for building advanced AI agents. It covers fundamental concepts such as multi-agent systems, memory management, planning, and reasoning loops. This resource is ideal for developers and researchers seeking practical insights into agentic AI development.