agentevals: Evaluate AI Agents from OpenTelemetry Traces

Summary
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 162
- Forks
- 28
- Added to OSRepos
- October 4, 2026
- Last analyzed
- October 4, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
agentevals evaluates agent behavior from recorded OpenTelemetry traces. It lets teams inspect and score what an agent actually did, such as which tools it called and what it answered, without rerunning the agent or spending tokens on another execution.
It is most useful for framework-agnostic, instrumented agents where behavior can be checked against expected tool trajectories, response criteria, or custom scoring rules. Evaluations can run offline from trace files, or through a local web interface and OTLP receiver. The project is under active development and notes that breaking changes are possible.
Key Features
- Reads Jaeger JSON and OTLP trace formats from OpenTelemetry-instrumented agents.
- Scores traces against golden evaluation sets, including tool-trajectory and response-match checks.
- Runs multiple evaluators and supports configurable trajectory matching.
- Supports custom evaluators in Python, JavaScript, or other languages, plus LLM-based judges.
- Provides an offline CLI for scripting and CI quality gates.
- Includes a web UI for trace inspection and live streaming through an OTLP receiver.
- Offers an MCP server for evaluating traces and inspecting live sessions.
- Can run locally without a required cloud service or database for basic evaluation.
Use Cases
- Agent developers can compare a recorded run with expected tool calls and responses while iterating on an instrumented agent.
- Platform and QA teams can add trace-based pass/fail checks to CI before deploying agent changes.
- Teams using multiple agent frameworks can reuse evaluation sets when each framework exports compatible OpenTelemetry traces.
- Operators debugging live behavior can stream traces to the UI and inspect tool calls and outputs as sessions run.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 162
- Forks: 28
- Topics: agentevals, agents, evals, evaluation, genai, llm, llm-as-judge, opentelemetry, otel, tracing
- Archived: false
Getting Started
Install the CLI, REST API, and embedded web UI from PyPI:
pip install agentevals-cli
For usage examples and optional integrations, see the repository README.
Alternatives
- opik: Opik is a broader tracing and evaluation platform, while agentevals focuses on scoring existing OpenTelemetry traces locally without rerunning agents.
- phoenix: Phoenix combines trace inspection with evaluation and troubleshooting in a self-hosted platform; agentevals focuses on local evaluation of existing traces.
- weave: Weave organizes and evaluates runs within an experiment-tracking workflow, while agentevals scores existing OpenTelemetry traces without extra model calls.
- langwatch: LangWatch adds simulation, prompt management, and production monitoring, while agentevals focuses on evaluating existing traces locally.
Considerations
- Evaluation depends on traces being available in supported Jaeger JSON or OTLP formats; frameworks need compatible OpenTelemetry instrumentation.
- Results reflect recorded behavior. The project is designed to score existing traces rather than execute agents or replay test cases.
- The README describes the project as under active development and warns that breaking changes may occur.
- Durable evaluation run history requires enabling the preview Postgres backend. Basic CLI evaluation and default in-memory server operation do not require a database.
Source repository
Open the original repository on GitHub.
Related repositories
Similar repositories that may be relevant next.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.

agent-anvil: Test AI Agent Tool Use in CI
October 1, 2026
Agent Anvil evaluates tool-using AI agents through scenario-based runs, trace checks, and optional semantic grading. It helps teams catch unsafe or incorrect tool behavior before it reaches production.