judgy: Estimate LLM Judge Success Rates with Bias Correction

judgy: Estimate LLM Judge Success Rates with Bias Correction

Summary

judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

At a glance

Language
Python
License
MIT
Stars
99
Forks
18
Added to OSRepos
December 7, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

judgy is a Python library for estimating the true success rate of a system evaluated by an LLM judge. It uses human labels on a test set to estimate the judge’s true positive and true negative rates, then adjusts the judge’s predictions on unlabeled data for those errors.

This is useful when manually labeling every example is impractical, but a representative human-labeled set is available to calibrate the judge. The estimate depends on the judge performing better than random chance, and the library also reports uncertainty through bootstrap confidence intervals.

Key Features

  • Corrects an observed pass rate using the judge’s estimated true positive and true negative rates.
  • Accepts binary human labels and judge predictions for a test set, plus judge predictions for unlabeled data.
  • Returns a corrected success-rate estimate and lower and upper confidence bounds.
  • Uses bootstrap resampling to estimate confidence intervals.
  • Allows the number of bootstrap iterations and confidence level to be configured.
  • Provides a Python API and can be installed from PyPI.

Use Cases

  • Evaluation teams estimating the pass rate of a model or system across many examples without collecting human labels for all of them.
  • Researchers calibrating an LLM judge against a human-labeled test set before using it to evaluate a larger dataset.
  • Developers tracking binary outcomes, such as good/bad or pass/fail, while accounting for known judge errors.

Project Facts

  • Language: Python
  • License: MIT
  • Stars: 99
  • Forks: 18
  • Archived: No

Getting Started

Install the package:

pip install judgy

See the README for the API example and further details.

Alternatives

  • promptbench: PromptBench evaluates model behavior across datasets, prompts, and attacks; judgy estimates pass rates while correcting for judge errors.
  • LLMBox: LLMBox provides broad model training and benchmark evaluation pipelines, while judgy focuses on calibrated pass-rate estimates from judge predictions.
  • AuditNLG: AuditNLG checks generated text for factualness, safety, and compliance; judgy corrects evaluation pass rates using human-labeled calibration data.
  • lighteval: Lighteval runs benchmark tasks across models and backends, while judgy estimates pass rates on unlabeled data using calibrated judge predictions.

Considerations

  • Requires Python 3.8 or later and NumPy 1.20.0 or later.
  • Requires human labels and judge predictions on a test set that can be used to estimate judge accuracy.
  • The correction assumes the judge performs better than random chance, expressed as TPR + TNR > 1. The README warns that the method may not apply when judge accuracy is too low.
  • Estimates are for binary outcomes, represented as 0/1 labels and predictions.

Comparisons

Source repository

Open the original repository on GitHub.

19 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️