Judgy: Correcting LLM Judge Bias for Reliable AI Model Evaluation
This repository profile is provided by osrepos.com, an open source repository discovery platform.

Summary
Judgy is a Python package designed to improve the reliability of evaluations performed by LLM-as-Judges. It provides tools to estimate the true success rate of a system by correcting for LLM judge bias and generating confidence intervals through bootstrapping. This helps ensure more accurate and trustworthy assessments of AI model performance.
Repository Information
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Introduction
judgy is a Python library developed to address the challenges of using Large Language Models (LLMs) as judges for evaluating other AI models or systems. When LLMs are used in this capacity, their inherent biases and potential errors can significantly impact the reliability and accuracy of the evaluation results. This package provides a robust solution to estimate the true success rate of your system by correcting for these LLM judge biases.
Installation
Getting started with judgy is straightforward. You can install the package using pip:
pip install judgy
Examples
Here's a quick example demonstrating how to use judgy to estimate a system's true success rate with confidence intervals:
import numpy as np
from judgy import estimate_success_rate
# Your data: 1 = Pass, 0 = Fail
test_labels = [1, 1, 0, 0, 1, 0, 1, 0] # Human labels on test set
test_preds = [1, 0, 0, 1, 1, 0, 1, 0] # LLM judge predictions on test set
unlabeled_preds = [1, 1, 0, 1, 0, 1, 0, 1] # LLM judge predictions on unlabeled data
# Estimate true pass rate with 95% confidence interval
theta_hat, lower_bound, upper_bound = estimate_success_rate(
test_labels=test_labels,
test_preds=test_preds,
unlabeled_preds=unlabeled_preds
)
print(f"Estimated true pass rate: {theta_hat:.3f}")
print(f"95% Confidence interval: [{lower_bound:.3f}, {upper_bound:.3f}]")
Why Use It
The core problem judgy solves is the unreliability introduced by LLM judge biases. By implementing a bias correction method, judgy helps you obtain a more accurate estimate of your system's true performance. It works by first estimating the LLM judge's accuracy (True Positive Rate and True Negative Rate) using a labeled test set. Then, it applies a correction formula to the observed pass rate from the judge. Finally, it uses bootstrap resampling to quantify the uncertainty and provide a confidence interval, giving you a statistically sound range for your system's true success rate. This ensures that your evaluations are not only corrected for bias but also come with a measure of their statistical robustness.
Links
For more detailed information, documentation, and to contribute, visit the official judgy resources:
Related repositories
Similar repositories that may be relevant next.

dify-official-plugins: Extending Dify with AI Models, Tools, and Agent Strategies
August 18, 2026
The `dify-official-plugins` repository hosts a collection of official plugins for Dify, an open-source platform for developing LLM-powered AI applications. These plugins, including models, tools, agent strategies, and extensions, enhance Dify's capabilities and are maintained by the official Dify team. They are designed to help developers efficiently build, deploy, and manage AI-driven solutions.

Agent Skills: A Standardized Way to Give AI Agents New Capabilities
August 18, 2026
Agent Skills provides a lightweight, open format for extending AI agent capabilities with specialized knowledge and workflows. It allows packaging procedural knowledge and context into portable, version-controlled folders that agents load on demand. This enables agents to gain domain expertise, follow repeatable workflows, and reuse skills across various compatible AI tools.

A-MEM: Self-Evolving Memory for Coding Agents
August 17, 2026
A-MEM is an innovative self-evolving memory system designed for coding agents, organizing knowledge into a dynamic Zettelkasten-style graph. It allows memories to evolve and connect over time, enhancing an agent's ability to recall and utilize information effectively. This system offers both semantic and structural search capabilities for a richer knowledge base.

Agent Sandbox: Secure Local Development for AI Coding Agents
August 17, 2026
Agent Sandbox provides a robust and secure local development environment specifically designed for collaborating with AI coding agents. It ensures minimal filesystem access, configurable network egress policies, and secure secret injection, protecting your local machine from potentially risky agent operations. This project supports various AI agents and integrates seamlessly with both CLI and popular IDE devcontainer setups.
Source repository
Open the original repository on GitHub.
17 counted GitHub visits