# Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Source: osrepos.com
Repository profile: https://osrepos.com/repo/agent-axiom-agent-anvil
Generated for open source discovery and AI-assisted research.

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

GitHub: https://github.com/agent-axiom/agent-anvil
OSRepos URL: https://osrepos.com/repo/agent-axiom-agent-anvil

## Summary

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

## Topics

- Python
- AI
- Agents
- Evaluation
- CI/CD
- Testing
- LLM
- Tools

## Repository Information

Last analyzed by OSRepos: Thu Oct 01 2026 17:41:55 GMT+0100 (Western European Summer Time)
Detail views: 1
GitHub clicks: 0

## Safety Notice

OSRepos shares public repositories for knowledge and discovery only. Review source code, dependencies, licenses, and security implications before running or installing anything.

## Content

## Introduction

Agent Anvil is an innovative, CI-first evaluation harness specifically built for tool-using AI agents. Unlike traditional evaluations that only assess the final answer, Agent Anvil focuses on the agent's behavior throughout the process. It aims to answer the critical question: "did the agent behave safely while getting there?"

## Why Use Agent Anvil & Key Benefits

Agent Anvil catches workflow bugs that final-answer evaluations often miss, such as incorrect tool usage, premature calls to destructive tools, missing clarifying questions, loops, and violations of business invariants. It achieves this by running YAML scenario suites, recording comprehensive model and tool traces, and performing both deterministic checks and OpenAI-powered semantic grading. Key benefits include:

*   **Behavioral Evaluation**: Focuses on the agent's process, not just the outcome, ensuring safe and compliant behavior.
*   **Comprehensive Trace Analysis**: Records detailed model and tool interactions, allowing for deep inspection of agent decisions.
*   **Failure Clustering & Repair Plans**: Automatically groups similar failures and generates concrete suggestions for fixing prompts, tool descriptions, and guardrails.
*   **CI/CD Integration**: Designed to integrate seamlessly into continuous integration pipelines, failing builds on regressions and providing actionable reports.
*   **Deterministic Assertions**: Supports a powerful DSL for defining trace invariants, including ordered tool calls, max call counts, forbidden argument values, and output checks.
*   **Leaderboard Submissions**: Facilitates verifiable public leaderboard submissions, offering transparency and provenance for agent performance.

## Installation

To get started with Agent Anvil, you can install its dependencies using `uv`:

bash
uv sync --group dev


## Examples

Agent Anvil provides a straightforward command-line interface for various tasks. Here are a few examples to illustrate its capabilities:

**Run a deterministic demo without OpenAI credentials:**

bash
uv run anvil run scenarios/external_jsonl_agent.yaml --offline


**Check conformance for an external agent:**

bash
uv run anvil conformance external-agent --agent-command "python my_agent.py"


**Run the intentional regression demo and inspect the repair plan:**

bash
uv run anvil run scenarios/refund_agent.yaml --offline --agent-mode offline --trials 1 || true
uv run anvil repair runs/latest


Agent Anvil's core loop is intentionally small and CI-shaped: `run -> trace -> check -> grade -> report -> CI`. You can see a visual demonstration of Agent Anvil catching a premature tool call in the project's README.

## Links

*   **GitHub Repository**: [agent-axiom/agent-anvil](https://github.com/agent-axiom/agent-anvil){:target="_blank"}
*   **3-minute judges guide**: [docs/judges-guide.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/judges-guide.md){:target="_blank"}
*   **Sample report**: [docs/demo-report.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/demo-report.md){:target="_blank"}
*   **Sample repair plan**: [docs/demo-repair-plan.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/demo-repair-plan.md){:target="_blank"}
*   **Leaderboard Submissions Guide**: [docs/leaderboard.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/leaderboard.md){:target="_blank"}
*   **CLI Reference**: [docs/cli.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/cli.md){:target="_blank"}