judgy vs lighteval
Python tools for evaluating model and system performance
judgy estimates a system’s binary success rate by calibrating LLM judge predictions against human labels. lighteval runs language model evaluations across tasks and inference backends, with sample-level results and options for custom tasks and metrics.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

lighteval: Evaluate Language Models Across Backends
Lighteval is a Python toolkit for running LLM evaluations across local models and remote inference backends. It combines a broad task catalog with custom metrics and detailed sample-level results for teams comparing or debugging model performance.
| judgy | lighteval | |
|---|---|---|
| Language | Python | Python |
| License | MIT | MIT |
| Stars | 99 | 2.5k |
| Forks | 18 | 566 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- judgy focuses on correcting pass-rate estimates for judge errors, while lighteval provides a toolkit for running language model evaluations.
- judgy uses human-labeled calibration data and bootstrap resampling to produce an estimate with confidence bounds; lighteval includes a catalog of more than 1,000 evaluation tasks.
- judgy is designed for binary outcomes, while lighteval covers domains including knowledge, math, coding, multilingual evaluation, and language understanding.
- judgy accepts test-set labels and predictions plus predictions for unlabeled data; lighteval supports local models, in-memory models, and supported inference backends.
- Both are Python projects under the MIT license; lighteval reports more stars and forks in the supplied project data.
Choose judgy if you…
- need to estimate binary pass rates on larger unlabeled datasets using human-labeled calibration data.
- want to account for an LLM judge’s estimated errors and report bootstrap confidence bounds.
Choose lighteval if you…
- need to run language model evaluations across supported local or remote inference backends.
- want a broad task catalog, sample-level results, or the ability to add custom tasks and metrics.
- need to compare models across established benchmarks or inspect individual responses.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.