Open Source Evaluation Tools
Evaluation is the systematic measurement of how well a model or AI application performs against defined goals. It helps teams compare systems, detect regressions, assess answer quality and safety, and understand where outputs fall short. Tests may use fixed datasets, simulated scenarios, human judgments, or automated metrics, depending on the task and the risks involved. For language models and agents, evaluation can also examine reasoning steps, tool use, retrieved information, and consistency across repeated runs.
Open source tools in this area include benchmark runners, dataset and test-case generators, scoring frameworks, experiment trackers, and systems for inspecting traces or monitoring results. When choosing one, consider its maturity, license, maintenance activity, supported models and integrations, data requirements, and whether its metrics fit your use case. These tools are useful to researchers, developers, and teams building or operating AI systems.
14 repositories · updated October 3, 2026

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

Ragas: Supercharge Your LLM Application Evaluations
Ragas is an ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. It offers objective metrics, intelligent test generation, and data-driven insights to move beyond subjective assessments. This framework helps developers build feedback loops and continuously improve their LLM applications.

RAGChecker: A Fine-grained Framework for Diagnosing RAG Systems
RAGChecker is an advanced automatic evaluation framework developed by Amazon Science, specifically designed to assess and diagnose Retrieval-Augmented Generation (RAG) systems. It offers a comprehensive suite of metrics and tools for in-depth analysis of RAG performance. This framework empowers developers and researchers to thoroughly evaluate and enhance their RAG systems with precision.

DeepFabric: High-Quality Synthetic Data for Agentic AI Systems
DeepFabric is an open-source Python library designed to generate high-quality synthetic training data for language models and agent evaluations. It excels at creating domain-specific datasets that teach models to think, plan, and act effectively, including correct tool usage and adherence to schema structures. This comprehensive pipeline also integrates training and evaluation capabilities, ensuring robust model development.

lighteval: Evaluate Language Models Across Backends
Lighteval is a Python toolkit for running LLM evaluations across local models and remote inference backends. It combines a broad task catalog with custom metrics and detailed sample-level results for teams comparing or debugging model performance.

AgentEvals: Robust Evaluation Tools for LLM Agent Trajectories
AgentEvals is a powerful open-source package from LangChain designed to simplify the evaluation of agentic applications. It provides a collection of ready-made evaluators and utilities, with a particular focus on analyzing agent trajectories, the intermediate steps an agent takes to solve problems. This helps developers understand and improve the reliability and performance of their LLM agents.

phoenix: Observe and Evaluate AI Applications
Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

open_deep_research: Build Configurable Deep Research Agents
A Python research agent that searches the web, synthesizes findings, and produces reports using configurable language models and search tools. It suits developers who want to customize or deploy a research workflow with LangGraph.

LangWatch: The Platform for LLM Evaluations and AI Agent Testing
LangWatch is an open-source platform designed for end-to-end LLM evaluations and AI agent testing. It helps teams test, simulate, evaluate, and monitor LLM-powered agents both before release and in production. Built for robust regression testing, simulations, and production observability, LangWatch eliminates the need for custom tooling.

Promptfoo: LLM Evaluation and Red Teaming for AI Applications
Promptfoo is an open-source CLI and library designed for evaluating and red-teaming Large Language Model (LLM) applications. It enables developers to test prompts, agents, and RAGs, compare model performance, and secure AI apps through vulnerability scanning. With simple declarative configs and CI/CD integration, Promptfoo helps ship reliable and secure AI solutions.

Langsmith-sdk: Client SDK for LLM Debugging, Evaluation, and Monitoring
The Langsmith-sdk provides client SDKs for interacting with the LangSmith platform, enabling robust debugging, evaluation, and monitoring of language models and intelligent agents. It offers native integrations with both LangChain Python and LangChain JS, making it an essential tool for LLM application development.

LLMBox: A Comprehensive Python Library for LLM Training and Evaluation
LLMBox is a comprehensive Python library designed for implementing Large Language Models, offering a unified training pipeline and extensive model evaluation capabilities. It provides a one-stop solution for both training and utilizing LLMs, emphasizing flexibility and efficiency. Developers can leverage its diverse training strategies and blazingly fast inference for their LLM projects.