judgy vs judges
LLM evaluation projects compared
judgy and judges are Python projects for evaluating system or model performance with LLM-based judgments. judgy calibrates a judge against human labels to estimate a corrected binary success rate, while judges provides reusable evaluators for assessing model inputs and outputs across a range of criteria.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

judges: Evaluate LLM Outputs with Reusable AI Judges
Databricks judges is a Python library for evaluating language model outputs with reusable LLM-based classifiers and graders. Use its research-backed judges, combine evaluations with a jury, or build a custom judge for your task.
| judgy | judges | |
|---|---|---|
| Language | Python | Python |
| License | MIT | Apache-2.0 |
| Stars | 99 | 339 |
| Forks | 18 | 37 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- judgy estimates a corrected pass rate from judge predictions and human-labeled calibration data; judges evaluates inputs and outputs with classifiers and graders.
- judgy focuses on binary outcomes and reports confidence bounds using bootstrap resampling; judges can return boolean judgments or numerical and Likert-scale scores, with reasoning.
- judgy requires a human-labeled test set to estimate judge accuracy; judges offers supplied evaluators, custom judges, and an AutoJudge option built from labeled examples and feedback.
- judgy is licensed under MIT and is not archived; judges is licensed under Apache-2.0 and its repository is archived.
- judgy lists 99 stars and 18 forks; judges lists 339 stars and 37 forks.
Choose judgy if you…
- need a corrected binary success-rate estimate from judge predictions and human calibration labels.
- want bootstrap confidence intervals for an estimate on a larger unlabeled dataset.
- need a Python library focused on accounting for known judge errors.
Choose judges if you…
- need reusable evaluators for criteria such as correctness, hallucination, safety, or relevance.
- want to combine multiple judges with a jury or create a custom evaluator.
- need a CLI for single or batch evaluations using JSON input.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.