agentevals: Evaluate AI Agents from OpenTelemetry Traces

agentevals: Evaluate AI Agents from OpenTelemetry Traces

Summary

agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

At a glance

Language
Python
License
Apache-2.0
Stars
162
Forks
28
Added to OSRepos
October 4, 2026
Last analyzed
October 4, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

agentevals evaluates agent behavior from recorded OpenTelemetry traces. It lets teams inspect and score what an agent actually did, such as which tools it called and what it answered, without rerunning the agent or spending tokens on another execution.

It is most useful for framework-agnostic, instrumented agents where behavior can be checked against expected tool trajectories, response criteria, or custom scoring rules. Evaluations can run offline from trace files, or through a local web interface and OTLP receiver. The project is under active development and notes that breaking changes are possible.

Key Features

  • Reads Jaeger JSON and OTLP trace formats from OpenTelemetry-instrumented agents.
  • Scores traces against golden evaluation sets, including tool-trajectory and response-match checks.
  • Runs multiple evaluators and supports configurable trajectory matching.
  • Supports custom evaluators in Python, JavaScript, or other languages, plus LLM-based judges.
  • Provides an offline CLI for scripting and CI quality gates.
  • Includes a web UI for trace inspection and live streaming through an OTLP receiver.
  • Offers an MCP server for evaluating traces and inspecting live sessions.
  • Can run locally without a required cloud service or database for basic evaluation.

Use Cases

  • Agent developers can compare a recorded run with expected tool calls and responses while iterating on an instrumented agent.
  • Platform and QA teams can add trace-based pass/fail checks to CI before deploying agent changes.
  • Teams using multiple agent frameworks can reuse evaluation sets when each framework exports compatible OpenTelemetry traces.
  • Operators debugging live behavior can stream traces to the UI and inspect tool calls and outputs as sessions run.

Project Facts

  • Language: Python
  • License: Apache-2.0
  • Stars: 162
  • Forks: 28
  • Topics: agentevals, agents, evals, evaluation, genai, llm, llm-as-judge, opentelemetry, otel, tracing
  • Archived: false

Getting Started

Install the CLI, REST API, and embedded web UI from PyPI:

pip install agentevals-cli

For usage examples and optional integrations, see the repository README.

Alternatives

  • opik: Opik is a broader tracing and evaluation platform, while agentevals focuses on scoring existing OpenTelemetry traces locally without rerunning agents.
  • phoenix: Phoenix combines trace inspection with evaluation and troubleshooting in a self-hosted platform; agentevals focuses on local evaluation of existing traces.
  • weave: Weave organizes and evaluates runs within an experiment-tracking workflow, while agentevals scores existing OpenTelemetry traces without extra model calls.
  • langwatch: LangWatch adds simulation, prompt management, and production monitoring, while agentevals focuses on evaluating existing traces locally.

Considerations

  • Evaluation depends on traces being available in supported Jaeger JSON or OTLP formats; frameworks need compatible OpenTelemetry instrumentation.
  • Results reflect recorded behavior. The project is designed to score existing traces rather than execute agents or replay test cases.
  • The README describes the project as under active development and warns that breaking changes may occur.
  • Durable evaluation run history requires enabling the preview Postgres backend. Basic CLI evaluation and default in-memory server operation do not require a database.

Source repository

Open the original repository on GitHub.

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️