# agentevals: Evaluate AI Agents from OpenTelemetry Traces

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Source: osrepos.com
Repository profile: https://osrepos.com/repo/agentevals-dev-agentevals
Generated for open source discovery and AI-assisted research.

agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

GitHub: https://github.com/agentevals-dev/agentevals
OSRepos URL: https://osrepos.com/repo/agentevals-dev-agentevals

## Summary

agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

## Topics

- python
- ai-agents
- evaluation
- llm
- opentelemetry
- cli
- observability
- local-first

## Repository Information

Last analyzed by OSRepos: Sun Oct 04 2026 12:03:04 GMT+0100 (Western European Summer Time)
Detail views: 1
GitHub clicks: 0

## Safety Notice

OSRepos shares public repositories for knowledge and discovery only. Review source code, dependencies, licenses, and security implications before running or installing anything.

## Content

## Overview

[agentevals](https://github.com/agentevals-dev/agentevals) evaluates agent behavior from recorded OpenTelemetry traces. It lets teams inspect and score what an agent actually did, such as which tools it called and what it answered, without rerunning the agent or spending tokens on another execution.

It is most useful for framework-agnostic, instrumented agents where behavior can be checked against expected tool trajectories, response criteria, or custom scoring rules. Evaluations can run offline from trace files, or through a local web interface and OTLP receiver. The project is under active development and notes that breaking changes are possible.

## Key Features

- Reads Jaeger JSON and OTLP trace formats from OpenTelemetry-instrumented agents.
- Scores traces against golden evaluation sets, including tool-trajectory and response-match checks.
- Runs multiple evaluators and supports configurable trajectory matching.
- Supports custom evaluators in Python, JavaScript, or other languages, plus LLM-based judges.
- Provides an offline CLI for scripting and CI quality gates.
- Includes a web UI for trace inspection and live streaming through an OTLP receiver.
- Offers an MCP server for evaluating traces and inspecting live sessions.
- Can run locally without a required cloud service or database for basic evaluation.

## Use Cases

- **Agent developers** can compare a recorded run with expected tool calls and responses while iterating on an instrumented agent.
- **Platform and QA teams** can add trace-based pass/fail checks to CI before deploying agent changes.
- **Teams using multiple agent frameworks** can reuse evaluation sets when each framework exports compatible OpenTelemetry traces.
- **Operators debugging live behavior** can stream traces to the UI and inspect tool calls and outputs as sessions run.

## Project Facts

- Language: Python
- License: Apache-2.0
- Stars: 162
- Forks: 28
- Topics: agentevals, agents, evals, evaluation, genai, llm, llm-as-judge, opentelemetry, otel, tracing
- Archived: false

## Getting Started

Install the CLI, REST API, and embedded web UI from PyPI:

```bash
pip install agentevals-cli
```

For usage examples and optional integrations, see the [repository README](https://github.com/agentevals-dev/agentevals#readme).

## Alternatives

- [opik](https://osrepos.com/repo/comet-ml-opik): Opik is a broader tracing and evaluation platform, while agentevals focuses on scoring existing OpenTelemetry traces locally without rerunning agents.
- [phoenix](https://osrepos.com/repo/arize-ai-phoenix): Phoenix combines trace inspection with evaluation and troubleshooting in a self-hosted platform; agentevals focuses on local evaluation of existing traces.
- [weave](https://osrepos.com/repo/wandb-weave): Weave organizes and evaluates runs within an experiment-tracking workflow, while agentevals scores existing OpenTelemetry traces without extra model calls.
- [langwatch](https://osrepos.com/repo/langwatch-langwatch): LangWatch adds simulation, prompt management, and production monitoring, while agentevals focuses on evaluating existing traces locally.

## Considerations

- Evaluation depends on traces being available in supported Jaeger JSON or OTLP formats; frameworks need compatible OpenTelemetry instrumentation.
- Results reflect recorded behavior. The project is designed to score existing traces rather than execute agents or replay test cases.
- The README describes the project as under active development and warns that breaking changes may occur.
- Durable evaluation run history requires enabling the preview Postgres backend. Basic CLI evaluation and default in-memory server operation do not require a database.