Open Source Evaluation Tools

Discover 55 open source Evaluation repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Evaluation projects here are most often combined with Python, LLM and Machine Learning. Last updated October 4, 2026.

55 repositories · updated October 4, 2026

sumy: Summarize Text and HTML Documents

sumy: Summarize Text and HTML Documents

sumy is a Python library and command-line tool that extracts concise summaries from plain text and HTML. It offers several extractive summarization methods, language tokenization support, and tools for evaluating summaries.

PythonNLPNatural Language Processing
Added Dec 14, 2025 View details
opik: Trace, Evaluate, and Monitor LLM Applications

opik: Trace, Evaluate, and Monitor LLM Applications

Opik is a platform for tracing and evaluating LLM applications, RAG systems, and AI agents. Teams can use it to inspect workflows, run evaluations, and monitor deployments, either self-hosted or through Comet Cloud.

PythonLLMObservability
Added Dec 13, 2025 View details
Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression

Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression

PCToolkit is a Python toolkit for applying and evaluating prompt-compression methods for large language models. It brings five compressors, datasets, and evaluation metrics behind modular interfaces, making it useful for comparing methods across language tasks.

PythonAILLM
Added Dec 13, 2025 View details
judgy: Estimate LLM Judge Success Rates with Bias Correction

judgy: Estimate LLM Judge Success Rates with Bias Correction

judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

PythonAILLM
Added Dec 7, 2025 View details
open-r1: Reproduce DeepSeek-R1 Training and Evaluation

open-r1: Reproduce DeepSeek-R1 Training and Evaluation

Open R1 is Hugging Face’s toolkit and research project for reproducing the DeepSeek-R1 pipeline with open datasets, training scripts, and evaluation workflows. It is no longer maintained; its training work has moved to TRL.

PythonMachine LearningLLM
Added Nov 17, 2025 View details
weave: Trace and Evaluate Generative AI Applications

weave: Trace and Evaluate Generative AI Applications

Weave helps developers trace language model inputs, outputs, and function calls, then organize and evaluate application runs. It suits teams building AI features that need debugging and repeatable comparisons across experiments and production.

PythonAILLM
Added Nov 3, 2025 View details
ai-samples: Learn AI Development Patterns for .NET

ai-samples: Learn AI Development Patterns for .NET

.NET code samples show how to integrate AI services, build chat experiences, call functions, and evaluate model responses. The repository is archived, so use it as a reference and follow the current .NET AI documentation for maintained guidance.

CsharpAIGenerative AI
Added Oct 11, 2025 View details
Previous Page 5 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️