agentevals vs skill-up
Agent evaluation tools compared
agentevals evaluates recorded OpenTelemetry traces, while skill-up runs repeatable test cases against agents, Skills, and workspaces. The main distinction is trace-based evaluation of past behavior versus executing configured cases and reporting their results.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

skill-up: Evaluate and Improve Agent Skills
skill-up is a Go CLI for testing Agent Skills, agents, and workspaces with repeatable cases and structured reports. It suits teams that want to catch behavior regressions or iteratively improve Skills using evaluation results.
| agentevals | skill-up | |
|---|---|---|
| Language | Python | Go |
| License | Apache-2.0 | Apache-2.0 |
| Stars | 162 | 1.1k |
| Forks | 28 | 95 |
| Last analyzed | Oct 4, 2026 | Oct 7, 2026 |
Key differences
- agentevals reads Jaeger JSON and OTLP traces; skill-up runs YAML-defined cases against supported agent engines.
- agentevals focuses on recorded tool trajectories and responses; skill-up also evaluates workspace changes and can compare runs with and without a Skill.
- agentevals is written in Python and offers a CLI, web UI, OTLP receiver, and MCP server; skill-up is a Go CLI with JSON, JUnit XML, HTML, and benchmark reports.
- agentevals can evaluate locally from traces without rerunning agents; skill-up requires an agent engine for evaluation and supports built-in or custom engines.
- Both are Apache-2.0 licensed. agentevals warns that breaking changes may occur, while skill-up is described as new and changing quickly.
Choose agentevals if you…
- already have compatible OpenTelemetry traces and want to score recorded behavior without rerunning agents.
- need offline trace evaluation, CI checks, or live trace inspection through a local UI.
- want to define custom evaluators or reuse evaluation sets across frameworks that export compatible traces.
Choose skill-up if you…
- want repeatable YAML test cases for Skills, agents, or workspace tasks.
- need to compare Skill-enabled runs with runs that omit the Skill, or test multiple agent engines.
- want structured reports for local review or CI, or an iteration loop using skill-upper.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.