Open Source LLM Projects
Large language models (LLMs) process and generate text, code, and other forms of content from prompts. They can help with tasks such as answering questions, summarizing documents, translating, and assisting with software development. Open source LLM technology gives people ways to inspect, adapt, and run models, while local inference and hosting options can help address privacy, cost, and connectivity needs.
This area includes model weights and training resources, inference engines, APIs, evaluation tools, and applications that connect models to data or other capabilities. When choosing a tool, check its license and usage terms, supported hardware, resource requirements, maintenance activity, and compatibility with your existing workflows. These tools are useful to developers, researchers, organizations, and individuals who want to build, study, or run language-model systems with greater control.
274 repositories · updated October 4, 2026

agentevals: Evaluate AI Agents from OpenTelemetry Traces
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

OrcaReplay: Record and Replay AI Agent Runs
OrcaReplay records AI agent runs so developers can inspect what happened, replay interactions offline, and fork a run from a checkpoint onto another model. It is a TypeScript CLI for debugging and comparing agent behavior without modifying the agent.

SwarmLLM: Run Local and Distributed AI Models
SwarmLLM runs open AI models on your computer and can pool resources with other computers to run larger models. It also provides OpenAI- and Anthropic-compatible APIs for local apps and agents.

openmake_llm: Coordinate AI Models, Agents, and Tools
OpenMake is a self-hosted AI workspace that coordinates chat models, agents, and tools in one place. It suits people who want multi-step AI workflows with local or open-weight models, while keeping infrastructure and model choices under their control.

benchmark-radar: Discover AI Benchmarks and Track Evaluation Scores
Benchmark Radar gathers AI benchmark records and evaluation evidence from public sources in a searchable catalog. Use it to discover benchmarks, inspect reported scores and citations, and follow new findings through its dashboard or offline CLI.

shimmy: Serve Local GGUF Models with an OpenAI-Compatible API
Shimmy is a Rust inference server that runs GGUF language models locally and exposes an OpenAI-compatible API. It suits developers who want a lightweight alternative for connecting existing tools to local models without Python or llama.cpp.

router: Route AI Requests to the Best Model
weave-os/router is a Go proxy that routes AI requests across configured model providers, while accepting Anthropic, OpenAI, and Gemini API formats. It suits developers who want model choice and routing behind one endpoint, including agent and coding-tool users.

fx: Run a Unix-Style Coding Agent in Your Terminal
fx is a native Zig coding-agent CLI with an interactive shell, one-shot requests, and embedding options. It suits developers who want an extensible, model-agnostic agent that fits terminal workflows rather than a full-screen IDE.

Graphon: Execute Agentic AI Workflows as Graphs
Graphon is a Python engine for building and running agentic AI workflows as graphs. It suits developers who need event-driven execution, shared workflow state, and extensible integrations rather than a standalone chat interface.

AgentAleph: Run a Local Coding Agent and Manage GGUF Models
Agent Aleph is a Linux desktop app for running local GGUF models and using them to inspect, edit, and build software projects. It combines model downloads and inference controls with an approval-based coding agent.

llm-d-router: Route Inference Requests Intelligently
llm-d Router directs inference requests using model-serving signals such as KV-cache locality, load, and priority. It is for teams running LLM serving on Kubernetes that need proxy-integrated routing and request flow control.

inference-gateway: Unify LLM Providers Behind One API
Inference Gateway proxies requests to cloud and local LLM providers through compatible APIs. It is suited to teams building provider-flexible AI services that need self-hosting, MCP tools, authentication, or observability.