Open Source Evaluation Tools
Discover 55 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 4, 2026.
55 repositories · updated October 4, 2026

sumy: Summarize Text and HTML Documents
sumy is a Python library and command-line tool that extracts concise summaries from plain text and HTML. It offers several extractive summarization methods, language tokenization support, and tools for evaluating summaries.

opik: Trace, Evaluate, and Monitor LLM Applications
Opik is a platform for tracing and evaluating LLM applications, RAG systems, and AI agents. Teams can use it to inspect workflows, run evaluations, and monitor deployments, either self-hosted or through Comet Cloud.

Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression
PCToolkit is a Python toolkit for applying and evaluating prompt-compression methods for large language models. It brings five compressors, datasets, and evaluation metrics behind modular interfaces, making it useful for comparing methods across language tasks.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

open-r1: Reproduce DeepSeek-R1 Training and Evaluation
Open R1 is Hugging Face’s toolkit and research project for reproducing the DeepSeek-R1 pipeline with open datasets, training scripts, and evaluation workflows. It is no longer maintained; its training work has moved to TRL.

weave: Trace and Evaluate Generative AI Applications
Weave helps developers trace language model inputs, outputs, and function calls, then organize and evaluate application runs. It suits teams building AI features that need debugging and repeatable comparisons across experiments and production.

ai-samples: Learn AI Development Patterns for .NET
.NET code samples show how to integrate AI services, build chat experiences, call functions, and evaluate model responses. The repository is archived, so use it as a reference and follow the current .NET AI documentation for maintained guidance.