Open Source Evaluation Tools

Evaluation is the systematic measurement of how well a model or AI application performs against defined goals. It helps teams compare systems, detect regressions, assess answer quality and safety, and understand where outputs fall short. Tests may use fixed datasets, simulated scenarios, human judgments, or automated metrics, depending on the task and the risks involved. For language models and agents, evaluation can also examine reasoning steps, tool use, retrieved information, and consistency across repeated runs.

Open source tools in this area include benchmark runners, dataset and test-case generators, scoring frameworks, experiment trackers, and systems for inspecting traces or monitoring results. When choosing one, consider its maturity, license, maintenance activity, supported models and integrations, data requirements, and whether its metrics fit your use case. These tools are useful to researchers, developers, and teams building or operating AI systems.

14 repositories · updated October 3, 2026

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents

Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.

PythonAIAgent
Added Oct 1, 2026 View details
Ragas: Supercharge Your LLM Application Evaluations

Ragas: Supercharge Your LLM Application Evaluations

Ragas is an ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. It offers objective metrics, intelligent test generation, and data-driven insights to move beyond subjective assessments. This framework helps developers build feedback loops and continuously improve their LLM applications.

EvaluationLLMLlmops
Added Aug 9, 2026 View details
RAGChecker: A Fine-grained Framework for Diagnosing RAG Systems

RAGChecker: A Fine-grained Framework for Diagnosing RAG Systems

RAGChecker is an advanced automatic evaluation framework developed by Amazon Science, specifically designed to assess and diagnose Retrieval-Augmented Generation (RAG) systems. It offers a comprehensive suite of metrics and tools for in-depth analysis of RAG performance. This framework empowers developers and researchers to thoroughly evaluate and enhance their RAG systems with precision.

PythonRAGLLM
Added Jul 4, 2026 View details
DeepFabric: High-Quality Synthetic Data for Agentic AI Systems

DeepFabric: High-Quality Synthetic Data for Agentic AI Systems

DeepFabric is an open-source Python library designed to generate high-quality synthetic training data for language models and agent evaluations. It excels at creating domain-specific datasets that teach models to think, plan, and act effectively, including correct tool usage and adherence to schema structures. This comprehensive pipeline also integrates training and evaluation capabilities, ensuring robust model development.

PythonAIMachine Learning
Added Jul 2, 2026 View details
lighteval: Evaluate Language Models Across Backends

lighteval: Evaluate Language Models Across Backends

Lighteval is a Python toolkit for running LLM evaluations across local models and remote inference backends. It combines a broad task catalog with custom metrics and detailed sample-level results for teams comparing or debugging model performance.

PythonLLMEvaluation
Added Jul 1, 2026 View details
AgentEvals: Robust Evaluation Tools for LLM Agent Trajectories

AgentEvals: Robust Evaluation Tools for LLM Agent Trajectories

AgentEvals is a powerful open-source package from LangChain designed to simplify the evaluation of agentic applications. It provides a collection of ready-made evaluators and utilities, with a particular focus on analyzing agent trajectories, the intermediate steps an agent takes to solve problems. This helps developers understand and improve the reliability and performance of their LLM agents.

PythonLLMAgent
Added Jun 30, 2026 View details
phoenix: Observe and Evaluate AI Applications

phoenix: Observe and Evaluate AI Applications

Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

PythonAILLM
Added Jun 28, 2026 View details
open_deep_research: Build Configurable Deep Research Agents

open_deep_research: Build Configurable Deep Research Agents

A Python research agent that searches the web, synthesizes findings, and produces reports using configurable language models and search tools. It suits developers who want to customize or deploy a research workflow with LangGraph.

PythonAI AgentsLLM
Added May 15, 2026 View details
LangWatch: The Platform for LLM Evaluations and AI Agent Testing

LangWatch: The Platform for LLM Evaluations and AI Agent Testing

LangWatch is an open-source platform designed for end-to-end LLM evaluations and AI agent testing. It helps teams test, simulate, evaluate, and monitor LLM-powered agents both before release and in production. Built for robust regression testing, simulations, and production observability, LangWatch eliminates the need for custom tooling.

AIAnalyticsDatasets
Added Apr 28, 2026 View details
Promptfoo: LLM Evaluation and Red Teaming for AI Applications

Promptfoo: LLM Evaluation and Red Teaming for AI Applications

Promptfoo is an open-source CLI and library designed for evaluating and red-teaming Large Language Model (LLM) applications. It enables developers to test prompts, agents, and RAGs, compare model performance, and secure AI apps through vulnerability scanning. With simple declarative configs and CI/CD integration, Promptfoo helps ship reliable and secure AI solutions.

LLMEvaluationRed Teaming
Added Mar 24, 2026 View details
Langsmith-sdk: Client SDK for LLM Debugging, Evaluation, and Monitoring

Langsmith-sdk: Client SDK for LLM Debugging, Evaluation, and Monitoring

The Langsmith-sdk provides client SDKs for interacting with the LangSmith platform, enabling robust debugging, evaluation, and monitoring of language models and intelligent agents. It offers native integrations with both LangChain Python and LangChain JS, making it an essential tool for LLM application development.

EvaluationLanguage ModelObservability
Added Mar 18, 2026 View details
LLMBox: A Comprehensive Python Library for LLM Training and Evaluation

LLMBox: A Comprehensive Python Library for LLM Training and Evaluation

LLMBox is a comprehensive Python library designed for implementing Large Language Models, offering a unified training pipeline and extensive model evaluation capabilities. It provides a one-stop solution for both training and utilizing LLMs, emphasizing flexibility and efficiency. Developers can leverage its diverse training strategies and blazingly fast inference for their LLM projects.

PythonLLMMachine Learning
Added Mar 16, 2026 View details
Previous Page 1 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️