Open Source Evaluation Tools

Discover 55 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 4, 2026.

55 repositories · updated October 4, 2026

LLMBox: Train and Evaluate Large Language Models

LLMBox: Train and Evaluate Large Language Models

LLMBox is a Python library for training and evaluating large language models through a unified pipeline. It suits researchers and developers who want configurable fine-tuning workflows and a broad set of model and benchmark evaluation options.

PythonLLMMachine Learning
Added Mar 16, 2026 View details
inspector: Test and Debug MCP Servers and Apps

inspector: Test and Debug MCP Servers and Apps

MCPJam Inspector helps developers test, inspect, and evaluate MCP servers and apps across clients and models. Use it to debug protocol and OAuth behavior, run repeatable evaluations, and check for regressions in CI.

TypeScriptMCPTesting
Added Mar 5, 2026 View details
helicone: Monitor and Route LLM Requests

helicone: Monitor and Route LLM Requests

Helicone is an LLM observability platform and AI gateway for teams building AI applications. It helps developers inspect requests, track costs and latency, evaluate outputs, and route traffic across model providers.

TypeScriptLLMObservability
Added Mar 1, 2026 View details
MonoPCC: Estimate Monocular Depth in Endoscopic Images

MonoPCC: Estimate Monocular Depth in Endoscopic Images

MonoPCC is a PyTorch method for self-supervised monocular depth estimation on endoscopic images, using a photometric-invariant cycle constraint. It is aimed at researchers reproducing or extending depth and pose experiments on surgical-video datasets.

PythonMachine LearningDeep Learning
Added Feb 27, 2026 View details
courses: Learn Claude API and Prompting Techniques

courses: Learn Claude API and Prompting Techniques

Anthropic’s courses repository offers hands-on learning materials for using Claude through its API. It suits developers learning prompting, evaluations, and tool use, especially when they want guided examples before building Claude-powered workflows.

AIEducationJupyter Notebook
Added Feb 22, 2026 View details
verifiers: Build Environments for LLM Training and Evaluation

verifiers: Build Environments for LLM Training and Evaluation

verifiers is a Python library for building environments to train and evaluate large language models. It fits teams developing LLM reinforcement-learning workflows or reusable evaluations, especially those using Prime Intellect's training tools.

PythonLibraryMachine Learning
Added Feb 14, 2026 View details
LLMSanitize: Detect Contamination in NLP Data and LLMs

LLMSanitize: Detect Contamination in NLP Data and LLMs

LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.

PythonLibraryMachine Learning
Added Feb 9, 2026 View details
giskard-oss: Test and Red-Team LLM Agents

giskard-oss: Test and Red-Team LLM Agents

Giskard is a Python toolkit for evaluating agent behavior and probing AI systems for vulnerabilities. It suits teams building LLM agents or RAG applications that need repeatable checks, safety testing, and adversarial evaluation.

PythonLLMAI Agents
Added Jan 30, 2026 View details
EasyEdit: Edit Knowledge and Steer Large Language Models

EasyEdit: Edit Knowledge and Steer Large Language Models

EasyEdit is a framework for changing specific knowledge or behavior in large language models and evaluating the effects. It brings together multiple editing and inference-time steering methods for researchers and developers comparing approaches or testing targeted edits.

PythonLLMMachine Learning
Added Jan 26, 2026 View details
agentic_security: Scan LLMs and Agent Workflows for Vulnerabilities

agentic_security: Scan LLMs and Agent Workflows for Vulnerabilities

Agentic Security is a Python toolkit for probing LLMs and agent workflows with jailbreaks, fuzzing, and multimodal inputs. It fits developers and security teams who want to run configurable scans against an API and use the results in CI.

PythonSecurityLLM
Added Jan 4, 2026 View details
Paper2Code: Generate Code Repositories from ML Papers

Paper2Code: Generate Code Repositories from ML Papers

Paper2Code is a research project that uses specialized LLM agents to turn machine-learning papers into code repositories. It suits researchers and developers exploring paper reproduction, with setup options for OpenAI APIs or vLLM.

PythonMachine LearningAI Agents
Added Jan 1, 2026 View details
transformerlab-app: Train and Evaluate AI Models in One Workspace

transformerlab-app: Train and Evaluate AI Models in One Workspace

Transformer Lab is a research workspace for training, fine-tuning, running, and evaluating AI models on local machines or GPU clusters. It suits individual researchers who want a unified UI and teams that need to coordinate jobs across existing infrastructure.

PythonMachine LearningLLM
Added Dec 31, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️