Open Source Observability Projects
Observability is the practice of understanding a system’s internal state through the data it produces, including metrics, logs, and traces. It helps teams investigate outages, find performance bottlenecks, track service health, and understand how changes affect users. By bringing operational signals together, observability can make complex applications and infrastructure easier to troubleshoot and maintain.
Open source tools in this area include telemetry collectors, data stores, dashboards, tracing systems, and alerting platforms. When choosing one, consider its maturity, license, maintenance activity, deployment requirements, data retention and query capabilities, and compatibility with your existing stack. Observability tools are useful to developers, operations teams, and organizations running distributed services, cloud infrastructure, or data-intensive applications.
41 repositories · updated October 4, 2026

agentevals: Evaluate AI Agents from OpenTelemetry Traces
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

ai-observer: Monitor AI Coding Assistant Usage Locally
AI Observer is a self-hosted OpenTelemetry backend and dashboard for tracking usage across local AI coding assistants. It brings token and cost data, traces, logs, and metrics together in DuckDB, with both OTLP ingestion and file-based import or watching for supported tools.

agent-observability: Monitor AI Coding Agents Locally
A self-hosted OpenTelemetry stack for monitoring Claude Code and OpenAI Codex. It routes telemetry to Prometheus, Loki, and Tempo, then presents usage, performance, and activity in Grafana dashboards.

OrcaReplay: Record and Replay AI Agent Runs
OrcaReplay records AI agent runs so developers can inspect what happened, replay interactions offline, and fork a run from a checkpoint onto another model. It is a TypeScript CLI for debugging and comparing agent behavior without modifying the agent.

shepherd: Supervise Agents with Reversible Execution Traces
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.

inference-gateway: Unify LLM Providers Behind One API
Inference Gateway proxies requests to cloud and local LLM providers through compatible APIs. It is suited to teams building provider-flexible AI services that need self-hosting, MCP tools, authentication, or observability.

agentskit: Build AI Agents in JavaScript
AgentsKit is a TypeScript toolkit for building JavaScript AI agents from composable packages for model adapters, runtime, tools, memory, retrieval, and user interfaces. It suits developers who need more than a chat SDK and want to keep provider and component choices flexible.

sandbase-harness: Run Self-Hosted AI Agent Sessions
SandBase Harness is a local-first runtime for running AI agents with persistent sessions, governed tools, sandboxed execution, and audit trails. It suits teams that need to operate and inspect agents on their own infrastructure rather than build only around a model SDK.

agenttrail: Observe AI Coding Agents Locally
Agenttrail turns local file changes and supported coding-agent activity into project maps and a 3D task view. It helps developers follow progress and spot work that needs attention without managing agents or sending data to a cloud service.

ADR: Secure and Monitor Enterprise AI Agents
ADR is an enterprise security toolkit for discovering AI tools, collecting agent activity, benchmarking defenses, and detecting risky behavior. It is aimed at security teams evaluating or monitoring AI agents across employee endpoints and customer-facing systems.

evidently: Evaluate and Monitor ML and LLM Systems
Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.

phoenix: Observe and Evaluate AI Applications
Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.