agent-observability: Self-Hosted Observability for AI Coding Agents

Summary
agent-observability offers a robust, self-hosted OpenTelemetry stack designed for AI coding agents like Claude Code and OpenAI Codex. It ensures all telemetry data, including model requests, tool executions, and session activity, remains local within your environment. This comprehensive solution provides ready-made Grafana dashboards for deep insights into agent performance and usage.
At a glance
- Added to OSRepos
- October 3, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
agent-observability is a powerful, self-hosted observability stack tailored for AI coding agents such as Claude Code and OpenAI Codex. It leverages the OpenTelemetry (OTel) standard to collect metrics, logs, and traces emitted by these agents. The core benefit is that all your telemetry data, including sensitive information like model requests, prompts, and edits, stays entirely within your own environment, never leaving for third-party services.
The stack comprises essential observability components: the OpenTelemetry Collector, Prometheus for metrics, Loki for logs, Tempo for traces, and Grafana for visualization. It provides everything needed to stand up this stack, including configuration files and deployment manifests for Docker Compose, Kubernetes, and Helm.
Why Use It and Key Features
Choosing agent-observability offers significant advantages, particularly for privacy-conscious developers and teams:
- Complete Data Privacy: All telemetry data, from model requests to tool executions and session activity, is collected and stored locally. Nothing is sent to Anthropic, OpenAI, or any other third party, giving you full control and compliance.
- Comprehensive Telemetry: The stack captures a wide array of data points, including token usage, API-equivalent costs (for Claude Code), cache hit rates, code modifications, tool call breakdowns, response latency, and error counts.
- Ready-Made Grafana Dashboards: Gain immediate insights with pre-built, intuitive Grafana dashboards for both Claude Code and OpenAI Codex. These dashboards visualize:
- Claude Code: API-equivalent cost, token usage by type and model, productivity metrics (code edits, commits), tool usage, MCP server attribution, and performance (latency, errors).
- OpenAI Codex: Conversation and turn activity, token breakdown (input, output, cached, reasoning), tool calls and decisions, code-activity proxies (derived from shell commands), and performance (latency, transport errors).
- Full Observability Stack: Implements the complete OpenTelemetry signals pipeline:
- Metrics: Collected by OTel Collector and stored in Prometheus.
- Logs: Collected by OTel Collector and stored in Loki.
- Traces: Collected by OTel Collector and stored in Tempo, with derived metrics sent to Prometheus.
- All visualized in Grafana.
- Flexible Deployment: Supports various environments with deployment options for Docker Compose, Kubernetes (Kustomize), and Helm, making it adaptable to your existing infrastructure.
Installation
Getting agent-observability up and running is straightforward. The repository provides detailed instructions for each deployment method.
First, clone the repository:
git clone https://github.com/KB1SLN-Labs/agent-observability.git
cd agent-observability
From there, you can choose your preferred deployment method:
- Docker Compose: Ideal for quick starts on a single machine. Run
docker compose up -dafter cloning. - Kubernetes (Kustomize): For existing Kubernetes clusters without Helm. Use
kubectl apply -k k8s/. - Helm: For Kubernetes clusters with Helm, offering easy customization and upgrades. Use
helm upgrade --install claude-code ./helm.
Detailed configuration steps for Claude Code and OpenAI Codex, including environment variables and config.toml settings, are provided in the project's README to point your agents to the OTel Collector.
Examples
The agent-observability stack delivers rich visualizations through its Grafana dashboards. For instance, the Claude Code dashboard offers a real-time view of API-equivalent cost burn rate, token usage broken down by model and source, code edit acceptance rates, and tool usage patterns.

Similarly, the Codex dashboard provides insights into conversation activity, token distribution, tool success rates, and various performance metrics, helping you understand how your OpenAI Codex agent operates.
Links
- GitHub Repository: https://github.com/KB1SLN-Labs/agent-observability
Source repository
Open the original repository on GitHub.
Related repositories
Similar repositories that may be relevant next.

OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation
October 2, 2026
OrcaReplay introduces "time travel" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.

agenttrail: Real-time Observability for Your AI Coding Agents
September 3, 2026
agenttrail provides a local, open-source observability layer for AI coding agents. It transforms agent plans, tool calls, and file changes from tools like Claude Code, OpenAI Codex, and Cursor into a live, zoomable project map. This allows developers to monitor agent progress and activity in real time, ensuring they know what their agents are doing as it happens.

ADR: Secure and Monitor Enterprise AI Agents
August 14, 2026
ADR is an enterprise security toolkit for discovering AI tools, collecting agent activity, benchmarking defenses, and detecting risky behavior. It is aimed at security teams evaluating or monitoring AI agents across employee endpoints and customer-facing systems.

evidently: Evaluate and Monitor ML and LLM Systems
June 30, 2026
Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.