judgy vs AuditNLG
Python tools for evaluating LLMs and generated text
judgy estimates a system’s true binary success rate by calibrating LLM judge predictions against human labels. AuditNLG checks generated text for factualness, safety, and instruction compliance, with explanations and rewrite suggestions; the projects address different evaluation needs.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

AuditNLG: Check and Improve Trust in Generated Text
AuditNLG is a Python library for evaluating generated text for factualness, safety, and instruction compliance. It combines model- and API-based checks with explanations and rewrite suggestions for research and language-model application teams.
| judgy | AuditNLG | |
|---|---|---|
| Language | Python | Python |
| License | MIT | BSD-3-Clause |
| Stars | 99 | 103 |
| Forks | 18 | 12 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- judgy focuses on correcting pass-rate estimates for judge errors, while AuditNLG assesses factualness, safety, and instruction compliance in text.
- judgy requires human-labeled calibration data and judge predictions; AuditNLG offers checks through existing models and third-party services.
- judgy returns a corrected success-rate estimate with bootstrap confidence bounds; AuditNLG can provide scores, metadata, explanations, and candidate rewrites.
- judgy is licensed under MIT, while AuditNLG is licensed under BSD-3-Clause.
- judgy is suited to binary outcomes and assumes the calibrated judge performs better than chance; AuditNLG's results depend on the selected evaluators, models, services, and data.
Choose judgy if you…
- need to estimate a binary system pass rate from a larger unlabeled dataset.
- have human-labeled examples to calibrate an LLM judge and account for its errors.
- want a Python API that reports a corrected estimate with bootstrap confidence bounds.
Choose AuditNLG if you…
- need to evaluate generated text for factualness, safety, or instruction compliance.
- want explanations or candidate rewrites to investigate problematic outputs.
- need a shared workflow for checks using models or third-party services, including a command-line option.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.