Open Source LLM Projects
Large language models (LLMs) process and generate text, code, and other forms of content from prompts. They can help with tasks such as answering questions, summarizing documents, translating, and assisting with software development. Open source LLM technology gives people ways to inspect, adapt, and run models, while local inference and hosting options can help address privacy, cost, and connectivity needs.
This area includes model weights and training resources, inference engines, APIs, evaluation tools, and applications that connect models to data or other capabilities. When choosing a tool, check its license and usage terms, supported hardware, resource requirements, maintenance activity, and compatibility with your existing workflows. These tools are useful to developers, researchers, organizations, and individuals who want to build, study, or run language-model systems with greater control.
239 repositories · updated October 3, 2026

colibri: Run Large MoE Models on Local Hardware
colibri is a C inference engine that streams Mixture-of-Experts model weights from disk, so large models can run without fitting entirely in RAM or VRAM. It supports CPU-only use and optional GPU backends, with speed depending on hardware and storage.

Prime Agent: A Self-Improving RLM Agent for Coding and Autonomous Tasks
Prime Agent is an open-source, self-improving Recursive Language Model (RLM) agent designed for coding workflows and long-running autonomous tasks. It integrates a persistent Python control environment with a durable harness state, allowing useful context and reusable patterns to persist across sessions. Built in Rust, this project aims to enhance developer productivity through programmatic control and autonomous capabilities.

SwarmLLM: Run Local AI Models and Team Up for Giant Distributed Inference
SwarmLLM is a free, open-source application that allows you to run AI chat models directly on your own computer. It uniquely enables multiple computers to team up over the internet, collectively running models too large for a single machine. This platform offers an OpenAI and Anthropic-compatible API, all without requiring accounts or cryptocurrency.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

OpenCompany: Run a Hive Mind of AI Agents for Your Business
OpenCompany is a Rust-powered platform that enables a single operator to run an entire business using a "hive mind" of AI agents. It orchestrates specialized agents to handle various company functions, allowing the human operator to focus on vision and critical decisions. This innovative approach moves beyond traditional multi-agent systems, fostering concurrent deliberation and task completion.

AI Session Search: Ultra-Fast AI Agent Session Analysis
AI Session Search (aise) is an ultra-fast, Rust-powered tool designed for searching and analyzing local AI agent coding sessions. It seamlessly integrates and indexes nine different session formats, including those from Claude, Codex, Cursor, and Gemini CLI. This powerful utility enables developers to quickly recover context, track agent behavior, and efficiently manage their AI-generated code history.

Open ACE: Self-Hosted AI Coding Agent Workspace and Governance Platform
Open ACE is an open-source, self-hosted platform designed for managing AI coding agents within enterprise environments. It provides a unified workspace for various AI tools, enabling remote execution and robust governance features for API keys, costs, and compliance. This platform is ideal for organizations integrating AI coding agents into their development workflows, especially those requiring private deployment and centralized control.

Graphon: A Python Graph Execution Engine for Agentic AI Workflows
Graphon is an innovative Python-based graph execution engine designed for building agentic AI workflows. It provides a robust framework for orchestrating complex AI tasks, featuring event-driven execution, graph validation, and shared runtime state. This evolving repository already includes a functional engine, built-in nodes, and end-to-end examples for developers.

llm-d Router: Intelligent Orchestration for Multi-Phase LLM Inference
The `llm-d-router` is a sophisticated Go service designed as an intelligent entry point for large language model (LLM) inference requests. It orchestrates complex, multi-phase LLM inference pipelines across specialized worker pools, routing requests through an Inference Gateway to disaggregated vLLM workers. By exposing OpenAI-compatible APIs, it simplifies integration and leverages Kubernetes for scalable, efficient deployments.

LLM Wiki: Build a Self-Maintaining, Interlinked Knowledge Base with AI
LLM Wiki is a powerful cross-platform desktop application designed to transform your documents into an organized, interlinked knowledge base automatically. Unlike traditional RAG systems, it incrementally builds and maintains a persistent wiki from your sources, ensuring knowledge is compiled once and kept current. This innovative approach offers a dynamic and evolving personal knowledge management solution.

HarnessRouter: Unified Interface for AI Agent Harnesses
HarnessRouter Community Edition provides a self-hosted, Apache-2.0 licensed unified interface for various AI agent harnesses like Codex, Claude Code, and Hermes. It allows users to run multiple agents through a single API, offering features such as sessions, streaming, file handling, and cancellation. The project implements the open-standard Unified Harness Protocol (UHP), ensuring users maintain control over their keys and infrastructure.

Maka: A High-Performance Agent Workspace for AI Tasks
Apache Maka (Incubating) is a high-performance agent workspace designed to maintain a complete, append-only record of all agent actions. It focuses on measurable performance, local-first operation, and robust recovery mechanisms. This project provides a unified execution authority for desktop, TUI, and CLI clients, ensuring consistent agent behavior across platforms.