agent-anvil vs skill-up
Agent evaluation tools compared
agent-anvil and skill-up both run repeatable evaluations of AI agent behavior and support CI workflows. agent-anvil focuses on tool-use traces and workflow safety, while skill-up evaluates Agent Skills, agents, and workspace tasks across supported engines.

agent-anvil: Test AI Agent Tool Use in CI
Agent Anvil evaluates tool-using AI agents through scenario-based runs, trace checks, and optional semantic grading. It helps teams catch unsafe or incorrect tool behavior before it reaches production.

skill-up: Evaluate and Improve Agent Skills
skill-up is a Go CLI for testing Agent Skills, agents, and workspaces with repeatable cases and structured reports. It suits teams that want to catch behavior regressions or iteratively improve Skills using evaluation results.
| agent-anvil | skill-up | |
|---|---|---|
| Language | Python | Go |
| License | MIT | Apache-2.0 |
| Stars | 9 | 1.1k |
| Forks | 2 | 95 |
| Last analyzed | Oct 4, 2026 | Oct 7, 2026 |
Key differences
- agent-anvil checks tool-call order, limits, arguments, results, and policy requirements; skill-up tests responses and workspace changes, including comparisons with and without a Skill.
- agent-anvil is written in Python and specifies Python 3.12 or newer; skill-up is written in Go and requires Go 1.25 to build from source, with Node.js also listed as a runtime requirement.
- agent-anvil is MIT licensed; skill-up is Apache-2.0 licensed.
- agent-anvil supports demo agents and external agents through command or HTTP adapters; skill-up includes engines such as Claude Code, Codex, and Qoder CLI, as well as custom engines.
- agent-anvil has 9 stars and 2 forks in the supplied facts; skill-up has 1.1k stars and 95 forks, and its project notes describe rapid change and a short history.
Choose agent-anvil if you…
- need to verify tool-call sequences, arguments, prerequisites, or other workflow safety rules.
- want trace-based evaluations that can run offline without OpenAI credentials.
- are integrating an existing command-line or HTTP agent and want to assess its tool behavior.
Choose skill-up if you…
- want to test whether an Agent Skill improves results by comparing runs with and without it.
- need evaluations of repository tasks, fixtures, workspace changes, or multiple agent engines.
- want structured JSON, JUnit XML, HTML, or benchmark reports for review or CI.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.