{"name":"Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents","description":"Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.","github":"https://github.com/agent-axiom/agent-anvil","url":"https://osrepos.com/repo/agent-axiom-agent-anvil","source":"osrepos.com","sourceDescription":"This repository profile is provided by osrepos.com, an open source repository discovery platform.","repositoryProfile":"https://osrepos.com/repo/agent-axiom-agent-anvil","generatedFor":"open source discovery and AI-assisted research","markdown":"https://osrepos.com/repo/agent-axiom-agent-anvil.md","json":"https://osrepos.com/repo/agent-axiom-agent-anvil.json","topics":["Python","AI","Agents","Evaluation","CI/CD","Testing","LLM","Tools"],"keywords":["Python","AI","Agents","Evaluation","CI/CD","Testing","LLM","Tools"],"stars":null,"summary":"Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.","content":"## Introduction\n\nAgent Anvil is an innovative, CI-first evaluation harness specifically built for tool-using AI agents. Unlike traditional evaluations that only assess the final answer, Agent Anvil focuses on the agent's behavior throughout the process. It aims to answer the critical question: \"did the agent behave safely while getting there?\"\n\n## Why Use Agent Anvil & Key Benefits\n\nAgent Anvil catches workflow bugs that final-answer evaluations often miss, such as incorrect tool usage, premature calls to destructive tools, missing clarifying questions, loops, and violations of business invariants. It achieves this by running YAML scenario suites, recording comprehensive model and tool traces, and performing both deterministic checks and OpenAI-powered semantic grading. Key benefits include:\n\n*   **Behavioral Evaluation**: Focuses on the agent's process, not just the outcome, ensuring safe and compliant behavior.\n*   **Comprehensive Trace Analysis**: Records detailed model and tool interactions, allowing for deep inspection of agent decisions.\n*   **Failure Clustering & Repair Plans**: Automatically groups similar failures and generates concrete suggestions for fixing prompts, tool descriptions, and guardrails.\n*   **CI/CD Integration**: Designed to integrate seamlessly into continuous integration pipelines, failing builds on regressions and providing actionable reports.\n*   **Deterministic Assertions**: Supports a powerful DSL for defining trace invariants, including ordered tool calls, max call counts, forbidden argument values, and output checks.\n*   **Leaderboard Submissions**: Facilitates verifiable public leaderboard submissions, offering transparency and provenance for agent performance.\n\n## Installation\n\nTo get started with Agent Anvil, you can install its dependencies using `uv`:\n\nbash\nuv sync --group dev\n\n\n## Examples\n\nAgent Anvil provides a straightforward command-line interface for various tasks. Here are a few examples to illustrate its capabilities:\n\n**Run a deterministic demo without OpenAI credentials:**\n\nbash\nuv run anvil run scenarios/external_jsonl_agent.yaml --offline\n\n\n**Check conformance for an external agent:**\n\nbash\nuv run anvil conformance external-agent --agent-command \"python my_agent.py\"\n\n\n**Run the intentional regression demo and inspect the repair plan:**\n\nbash\nuv run anvil run scenarios/refund_agent.yaml --offline --agent-mode offline --trials 1 || true\nuv run anvil repair runs/latest\n\n\nAgent Anvil's core loop is intentionally small and CI-shaped: `run -> trace -> check -> grade -> report -> CI`. You can see a visual demonstration of Agent Anvil catching a premature tool call in the project's README.\n\n## Links\n\n*   **GitHub Repository**: [agent-axiom/agent-anvil](https://github.com/agent-axiom/agent-anvil){:target=\"_blank\"}\n*   **3-minute judges guide**: [docs/judges-guide.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/judges-guide.md){:target=\"_blank\"}\n*   **Sample report**: [docs/demo-report.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/demo-report.md){:target=\"_blank\"}\n*   **Sample repair plan**: [docs/demo-repair-plan.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/demo-repair-plan.md){:target=\"_blank\"}\n*   **Leaderboard Submissions Guide**: [docs/leaderboard.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/leaderboard.md){:target=\"_blank\"}\n*   **CLI Reference**: [docs/cli.md](https://github.com/agent-axiom/agent-anvil/blob/main/docs/cli.md){:target=\"_blank\"}","metrics":{"detailViews":1,"githubClicks":0},"dates":{"published":null,"modified":"2026-10-01T16:41:55.000Z"}}