Open Source Evaluation Tools

Discover 55 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 4, 2026.

55 repositories · updated October 4, 2026

phoenix: Observe and Evaluate AI Applications

phoenix: Observe and Evaluate AI Applications

Arize Phoenix is a self-hosted platform for tracing, evaluating, and troubleshooting LLM applications. It helps AI engineers inspect runtime behavior and test changes to prompts, models, and retrieval.

PythonAILLM
Added Jun 28, 2026 View details
EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks

EasyJailbreak: Build and Evaluate LLM Jailbreak Attacks

EasyJailbreak is a Python framework for assembling and testing jailbreak methods against language models. It suits researchers and developers who need reusable attack components and a structured way to evaluate model responses.

PythonLLMSecurity
Added Jun 26, 2026 View details
AuditNLG: Check and Improve Trust in Generated Text

AuditNLG: Check and Improve Trust in Generated Text

AuditNLG is a Python library for evaluating generated text for factualness, safety, and instruction compliance. It combines model- and API-based checks with explanations and rewrite suggestions for research and language-model application teams.

PythonLibraryAI
Added Jun 25, 2026 View details
MarkLLM: Implement and Evaluate LLM Watermarking

MarkLLM: Implement and Evaluate LLM Watermarking

MarkLLM is a Python toolkit for implementing, visualizing, and evaluating text watermarking methods for large language models. It helps researchers compare watermark detection, robustness, and text-quality effects through shared APIs and evaluation pipelines.

PythonLLMMachine Learning
Added Jun 23, 2026 View details
open_deep_research: Build Configurable Deep Research Agents

open_deep_research: Build Configurable Deep Research Agents

A Python research agent that searches the web, synthesizes findings, and produces reports using configurable language models and search tools. It suits developers who want to customize or deploy a research workflow with LangGraph.

PythonAI AgentsLLM
Added May 15, 2026 View details
llmgym: Build and Benchmark LLM Applications That Learn

llmgym: Build and Benchmark LLM Applications That Learn

LLMGym provides a shared environment interface for developing and evaluating LLM applications that learn from feedback. It is aimed at researchers and developers who want to test agents across varied interactive tasks.

PythonAILLM
Added May 12, 2026 View details
langwatch: Test and Monitor LLM Applications and AI Agents

langwatch: Test and Monitor LLM Applications and AI Agents

LangWatch helps teams evaluate, test, and monitor LLM applications and AI agents in production. It combines observability, simulation testing, prompt management, an AI gateway, and governance, with cloud and self-hosted options.

TypeScriptAILLM
Added Apr 28, 2026 View details
asta-paper-finder: Find Papers from Natural-Language Queries

asta-paper-finder: Find Papers from Natural-Language Queries

PaperFinder is a single-query research agent that finds papers using content and metadata criteria, then judges and ranks results. This frozen evaluation release is suited to reproducing results or running the core workflow locally, rather than replacing the live service.

PythonAI AgentsAI
Added Apr 24, 2026 View details
judges: Evaluate LLM Outputs with Reusable AI Judges

judges: Evaluate LLM Outputs with Reusable AI Judges

Databricks judges is a Python library for evaluating language model outputs with reusable LLM-based classifiers and graders. Use its research-backed judges, combine evaluations with a jury, or build a custom judge for your task.

PythonLLMEvaluation
Added Apr 14, 2026 View details
opik-mcp: Connect AI Assistants to Opik

opik-mcp: Connect AI Assistants to Opik

opik-mcp is a Python MCP server that lets AI coding assistants work with an Opik workspace. Use it to inspect traces, record scores, manage prompts, and explore evaluation data through natural-language workflows.

PythonMCPMCP Server
Added Mar 28, 2026 View details
promptfoo: Evaluate and Red-Team LLM Applications

promptfoo: Evaluate and Red-Team LLM Applications

Promptfoo is a CLI and library for evaluating prompts, comparing models, and testing LLM applications for security risks. It suits developers who want repeatable quality and vulnerability checks locally or in CI/CD.

TypeScriptLLMEvaluation
Added Mar 24, 2026 View details
langsmith-sdk: Trace and Evaluate LLM Applications

langsmith-sdk: Trace and Evaluate LLM Applications

LangSmith SDKs let Python and JavaScript/TypeScript applications send traces to the LangSmith platform for debugging, evaluation, and monitoring. Use them when you need visibility into model or agent behavior during development or operation.

PythonJavaScriptTypeScript
Added Mar 18, 2026 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️