Open Source Evaluation Tools
Discover 55 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 4, 2026.
55 repositories · updated October 4, 2026

LLMBox: Train and Evaluate Large Language Models
LLMBox is a Python library for training and evaluating large language models through a unified pipeline. It suits researchers and developers who want configurable fine-tuning workflows and a broad set of model and benchmark evaluation options.

inspector: Test and Debug MCP Servers and Apps
MCPJam Inspector helps developers test, inspect, and evaluate MCP servers and apps across clients and models. Use it to debug protocol and OAuth behavior, run repeatable evaluations, and check for regressions in CI.
helicone: Monitor and Route LLM Requests
Helicone is an LLM observability platform and AI gateway for teams building AI applications. It helps developers inspect requests, track costs and latency, evaluate outputs, and route traffic across model providers.

MonoPCC: Estimate Monocular Depth in Endoscopic Images
MonoPCC is a PyTorch method for self-supervised monocular depth estimation on endoscopic images, using a photometric-invariant cycle constraint. It is aimed at researchers reproducing or extending depth and pose experiments on surgical-video datasets.

courses: Learn Claude API and Prompting Techniques
Anthropic’s courses repository offers hands-on learning materials for using Claude through its API. It suits developers learning prompting, evaluations, and tool use, especially when they want guided examples before building Claude-powered workflows.

verifiers: Build Environments for LLM Training and Evaluation
verifiers is a Python library for building environments to train and evaluate large language models. It fits teams developing LLM reinforcement-learning workflows or reusable evaluations, especially those using Prime Intellect's training tools.

LLMSanitize: Detect Contamination in NLP Data and LLMs
LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.

giskard-oss: Test and Red-Team LLM Agents
Giskard is a Python toolkit for evaluating agent behavior and probing AI systems for vulnerabilities. It suits teams building LLM agents or RAG applications that need repeatable checks, safety testing, and adversarial evaluation.

EasyEdit: Edit Knowledge and Steer Large Language Models
EasyEdit is a framework for changing specific knowledge or behavior in large language models and evaluating the effects. It brings together multiple editing and inference-time steering methods for researchers and developers comparing approaches or testing targeted edits.

agentic_security: Scan LLMs and Agent Workflows for Vulnerabilities
Agentic Security is a Python toolkit for probing LLMs and agent workflows with jailbreaks, fuzzing, and multimodal inputs. It fits developers and security teams who want to run configurable scans against an API and use the results in CI.

Paper2Code: Generate Code Repositories from ML Papers
Paper2Code is a research project that uses specialized LLM agents to turn machine-learning papers into code repositories. It suits researchers and developers exploring paper reproduction, with setup options for OpenAI APIs or vLLM.

transformerlab-app: Train and Evaluate AI Models in One Workspace
Transformer Lab is a research workspace for training, fine-tuning, running, and evaluating AI models on local machines or GPU clusters. It suits individual researchers who want a unified UI and teams that need to coordinate jobs across existing infrastructure.