judgy vs LLMBox
Python tools for LLM evaluation compared
judgy estimates a system’s binary success rate by calibrating LLM judge predictions against human labels and correcting for judge errors. LLMBox is a broader library for training and evaluating language models, with fine-tuning workflows and benchmark evaluation options.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

LLMBox: Train and Evaluate Large Language Models
LLMBox is a Python library for training and evaluating large language models through a unified pipeline. It suits researchers and developers who want configurable fine-tuning workflows and a broad set of model and benchmark evaluation options.
| judgy | LLMBox | |
|---|---|---|
| Language | Python | Python |
| License | MIT | MIT |
| Stars | 99 | 848 |
| Forks | 18 | 104 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- judgy focuses on binary pass-rate estimation from a human-labeled calibration set; LLMBox covers model training and multiple evaluation workflows.
- judgy corrects judge predictions and reports bootstrap confidence intervals; LLMBox supports evaluation methods such as generation, likelihood-based ranking, and option probabilities.
- judgy is suited to estimating outcomes on unlabeled data after calibrating a judge; LLMBox is suited to fine-tuning models, preparing data, and running evaluations across supported benchmarks.
- Both projects are Python libraries with MIT licenses. LLMBox lists 848 stars and 104 forks, while judgy lists 99 stars and 18 forks.
- judgy requires binary labels and a judge that performs better than random chance; LLMBox workflows can require substantial compute and model-specific dependencies.
Choose judgy if you…
- need to estimate binary success rates on a larger unlabeled dataset using human-labeled calibration data.
- want to correct LLM judge errors and quantify uncertainty with bootstrap confidence intervals.
Choose LLMBox if you…
- need a unified library for fine-tuning and evaluating supported language models.
- want benchmark evaluation, data construction options, or training workflows such as supervised fine-tuning, PPO, or DPO.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.