LLM Evaluation

LLM evaluation is the practice of measuring how well large language models and the applications built around them perform. It helps teams compare models, test prompts, identify failures, and assess qualities such as accuracy, relevance, safety, and consistency. Because outputs can vary across inputs and runs, evaluation combines representative test cases, quantitative metrics, and human review to reveal problems that are difficult to detect through casual use alone.

Open source tools in this area include benchmarking suites, test and dataset generators, metric libraries, and platforms for tracing and reviewing model behavior. When choosing one, consider its evaluation methods, support for your models and workflows, data privacy, license, maintenance activity, and setup requirements. These tools are useful to researchers, developers, and teams building or maintaining LLM-powered products who need evidence to guide model selection and improve reliability.

4 repositories · updated October 2, 2026

OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation

OrcaReplay: Time Travel for AI Agents, Debugging and Evaluation

OrcaReplay introduces "time travel" capabilities for AI agents, allowing developers to record, replay, fork, and debug any agent run with any model. It addresses the challenges of AI agent debugging by providing byte-for-byte reproducibility, offline analysis, and the ability to compare different models from specific checkpoints. This tool, built by the OrcaRouter.ai team, enhances observability and control over complex agent behaviors.

Agent DebuggingAI AgentsLLM Agents
Added Oct 2, 2026 View details
Benchmark Radar: A Living Database for AI Benchmarks and Evaluation

Benchmark Radar: A Living Database for AI Benchmarks and Evaluation

Benchmark Radar is an extensive open-source project that tracks over 20,710 AI benchmark, evaluation, dataset, and data-quality records from 37 public sources. It provides daily updates, linked evidence, and tools for researchers and developers to discover and analyze AI benchmarks. This project is essential for anyone needing to stay current with AI evaluation trends and model performance.

AI BenchmarkLLM EvaluationAgentic Benchmarking
Added Sep 29, 2026 View details
Ragas: Supercharge Your LLM Application Evaluations

Ragas: Supercharge Your LLM Application Evaluations

Ragas is an ultimate toolkit for evaluating and optimizing Large Language Model (LLM) applications. It offers objective metrics, intelligent test generation, and data-driven insights to move beyond subjective assessments. This framework helps developers build feedback loops and continuously improve their LLM applications.

EvaluationLLMLlmops
Added Aug 9, 2026 View details
Opik: Open-Source LLM Observability, Evaluation, and Optimization

Opik: Open-Source LLM Observability, Evaluation, and Optimization

Opik is an open-source platform by Comet designed to streamline the lifecycle of LLM applications. It provides comprehensive tools for debugging, evaluating, and monitoring RAG systems and agentic workflows. Developers can leverage its tracing, automated evaluations, and production-ready dashboards to build and optimize generative AI applications.

LLM ObservabilityLLM EvaluationLlmops
Added Dec 13, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️