Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

This repository profile is provided by osrepos.com, an open source repository discovery platform.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Summary

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

Repository Information

Analyzed by OSRepos on October 1, 2026

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Introduction

Agent Anvil is an innovative, CI-first evaluation harness specifically built for tool-using AI agents. Unlike traditional evaluations that only assess the final answer, Agent Anvil focuses on the agent's behavior throughout the process. It aims to answer the critical question: "did the agent behave safely while getting there?"

Why Use Agent Anvil & Key Benefits

Agent Anvil catches workflow bugs that final-answer evaluations often miss, such as incorrect tool usage, premature calls to destructive tools, missing clarifying questions, loops, and violations of business invariants. It achieves this by running YAML scenario suites, recording comprehensive model and tool traces, and performing both deterministic checks and OpenAI-powered semantic grading. Key benefits include:

  • Behavioral Evaluation: Focuses on the agent's process, not just the outcome, ensuring safe and compliant behavior.
  • Comprehensive Trace Analysis: Records detailed model and tool interactions, allowing for deep inspection of agent decisions.
  • Failure Clustering & Repair Plans: Automatically groups similar failures and generates concrete suggestions for fixing prompts, tool descriptions, and guardrails.
  • CI/CD Integration: Designed to integrate seamlessly into continuous integration pipelines, failing builds on regressions and providing actionable reports.
  • Deterministic Assertions: Supports a powerful DSL for defining trace invariants, including ordered tool calls, max call counts, forbidden argument values, and output checks.
  • Leaderboard Submissions: Facilitates verifiable public leaderboard submissions, offering transparency and provenance for agent performance.

Installation

To get started with Agent Anvil, you can install its dependencies using uv:

uv sync --group dev

Examples

Agent Anvil provides a straightforward command-line interface for various tasks. Here are a few examples to illustrate its capabilities:

Run a deterministic demo without OpenAI credentials:

uv run anvil run scenarios/external_jsonl_agent.yaml --offline

Check conformance for an external agent:

uv run anvil conformance external-agent --agent-command "python my_agent.py"

Run the intentional regression demo and inspect the repair plan:

uv run anvil run scenarios/refund_agent.yaml --offline --agent-mode offline --trials 1 || true
uv run anvil repair runs/latest

Agent Anvil's core loop is intentionally small and CI-shaped: run -> trace -> check -> grade -> report -> CI. You can see a visual demonstration of Agent Anvil catching a premature tool call in the project's README.

Links

Related repositories

Similar repositories that may be relevant next.

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models

VeRL-Omni: Multimodal RL Training Framework for Diffusion & Omni Models

October 1, 2026

VeRL-Omni is a powerful RL training framework designed specifically for multimodal generative models, including diffusion models and omni-modality models. Built on top of the `verl` project, it offers easy, fast, and stable training solutions for complex generative AI tasks. The framework addresses unique challenges in multimodal RL, providing optimized performance and stability.

diffusion-modelsflow-matchingmultimodal
Uni-Agent: A Scalable Framework for Training Long-Horizon AI Agents

Uni-Agent: A Scalable Framework for Training Long-Horizon AI Agents

October 1, 2026

Uni-Agent is a powerful Python framework designed for training long-horizon agents at scale. It allows users to integrate existing agent harnesses, unify diverse agent tasks through an extensible interface, and run thousands of sessions concurrently for efficient data collection and training.

PythonReinforcement LearningAI Agents
Benchmark Radar: A Living Database for AI Benchmarks and Evaluation

Benchmark Radar: A Living Database for AI Benchmarks and Evaluation

September 29, 2026

Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.

AI BenchmarkLLM EvaluationAgentic Benchmarking
Pydantic AI Harness: Enhancing Your AI Agents with Robust Capabilities

Pydantic AI Harness: Enhancing Your AI Agents with Robust Capabilities

September 28, 2026

Pydantic AI Harness is the official capability and harness library for Pydantic AI, designed to extend agents for complex, long-running tasks. It provides a modular system of "capabilities" for functionalities like file system interaction, web research, memory, and sub-agent delegation. This library enables developers to build sophisticated and durable AI agents with ease.

PythonAIAgents

Source repository

Open the original repository on GitHub.

View on GitHub
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️