Open Source Evaluation Tools
Discover 43 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 3, 2026.
43 repositories · updated October 3, 2026

evalplus: Rigorously Evaluate LLM-Generated Code
EvalPlus evaluates code generated by language models with expanded correctness tests for HumanEval and MBPP, plus efficiency checks through EvalPerf. It is for researchers and developers comparing models or validating generated code more rigorously.

agentevals: Evaluate AI Agent Execution Trajectories
AgentEvals provides Python and TypeScript evaluators for checking the steps AI agents take, including tool calls and graph paths. Use it to compare runs with references or have an LLM judge trajectory quality.

evidently: Evaluate and Monitor ML and LLM Systems
Evidently is a Python framework for evaluating, testing, and monitoring machine-learning and LLM systems, from data quality to generated text. Use it to build offline reports and regression checks or track metrics over time in a monitoring dashboard.

phoenix: Observe and Evaluate AI Applications
Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks
EasyJailbreak is a Python framework for assembling and testing jailbreak methods against language models. It suits researchers and developers who need reusable attack components and a structured way to evaluate model responses.

AuditNLG: Check and Improve Trust in Generated Text
AuditNLG is a Python library for evaluating generated text for factualness, safety, and instruction compliance. It combines model- and API-based checks with explanations and rewrite suggestions for research and language-model application teams.

MarkLLM: Implement and Evaluate LLM Watermarking
MarkLLM is a Python toolkit for implementing, visualizing, and evaluating text watermarking methods for large language models. It helps researchers compare watermark detection, robustness, and text-quality effects through shared APIs and evaluation pipelines.

open_deep_research: Build Configurable Deep Research Agents
A Python research agent that searches the web, synthesizes findings, and produces reports using configurable language models and search tools. It suits developers who want to customize or deploy a research workflow with LangGraph.

llmgym: Build and Benchmark LLM Applications That Learn
LLMGym provides a shared environment interface for developing and evaluating LLM applications that learn from feedback. It is aimed at researchers and developers who want to test agents across varied interactive tasks.

langwatch: Test and Monitor LLM Applications and AI Agents
LangWatch helps teams evaluate, test, and monitor LLM applications and AI agents in production. It combines observability, simulation testing, prompt management, an AI gateway, and governance, with cloud and self-hosted options.

asta-paper-finder: Find Papers from Natural-Language Queries
PaperFinder is a single-query research agent that finds papers using content and metadata criteria, then judges and ranks results. This frozen evaluation release is suited to reproducing results or running the core workflow locally, rather than replacing the live service.

opik-mcp: Connect AI Assistants to Opik
opik-mcp is a Python MCP server that lets AI coding assistants work with an Opik workspace. Use it to inspect traces, record scores, manage prompts, and explore evaluation data through natural-language workflows.