judgy: Estimate LLM Judge Success Rates with Bias Correction

Summary
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 99
- Forks
- 18
- Added to OSRepos
- December 7, 2025
- Last analyzed
- October 3, 2026
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
judgy is a Python library for estimating the true success rate of a system evaluated by an LLM judge. It uses human labels on a test set to estimate the judge’s true positive and true negative rates, then adjusts the judge’s predictions on unlabeled data for those errors.
This is useful when manually labeling every example is impractical, but a representative human-labeled set is available to calibrate the judge. The estimate depends on the judge performing better than random chance, and the library also reports uncertainty through bootstrap confidence intervals.
Key Features
- Corrects an observed pass rate using the judge’s estimated true positive and true negative rates.
- Accepts binary human labels and judge predictions for a test set, plus judge predictions for unlabeled data.
- Returns a corrected success-rate estimate and lower and upper confidence bounds.
- Uses bootstrap resampling to estimate confidence intervals.
- Allows the number of bootstrap iterations and confidence level to be configured.
- Provides a Python API and can be installed from PyPI.
Use Cases
- Evaluation teams estimating the pass rate of a model or system across many examples without collecting human labels for all of them.
- Researchers calibrating an LLM judge against a human-labeled test set before using it to evaluate a larger dataset.
- Developers tracking binary outcomes, such as good/bad or pass/fail, while accounting for known judge errors.
Project Facts
- Language: Python
- License: MIT
- Stars: 99
- Forks: 18
- Archived: No
Getting Started
Install the package:
pip install judgy
See the README for the API example and further details.
Alternatives
- promptbench: PromptBench evaluates model behavior across datasets, prompts, and attacks; judgy estimates pass rates while correcting for judge errors.
- LLMBox: LLMBox provides broad model training and benchmark evaluation pipelines, while judgy focuses on calibrated pass-rate estimates from judge predictions.
- AuditNLG: AuditNLG checks generated text for factualness, safety, and compliance; judgy corrects evaluation pass rates using human-labeled calibration data.
- lighteval: Lighteval runs benchmark tasks across models and backends, while judgy estimates pass rates on unlabeled data using calibrated judge predictions.
Considerations
- Requires Python 3.8 or later and NumPy 1.20.0 or later.
- Requires human labels and judge predictions on a test set that can be used to estimate judge accuracy.
- The correction assumes the judge performs better than random chance, expressed as TPR + TNR > 1. The README warns that the method may not apply when judge accuracy is too low.
- Estimates are for binary outcomes, represented as 0/1 labels and predictions.
Comparisons
Source repository
Open the original repository on GitHub.
19 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
October 4, 2026
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.