judgy vs promptbench
LLM evaluation tools compared
judgy estimates binary success rates by correcting LLM judge predictions against human-labeled calibration data. promptbench provides broader evaluation of language and multimodal models, including prompt techniques, adversarial attacks, and dynamic evaluation.

judgy: Estimate LLM Judge Success Rates with Bias Correction
judgy estimates a system’s true pass rate from human-labeled calibration data and LLM judge predictions. It corrects for judge errors and uses bootstrap resampling to produce a confidence interval, making it useful when evaluating larger unlabeled datasets.

promptbench: Evaluate LLMs and Test Prompt Robustness
PromptBench is a Python library for evaluating language and multimodal models across datasets, prompting methods, and adversarial attacks. It suits researchers and developers comparing model behavior or studying robustness and dynamic evaluation.
| judgy | promptbench | |
|---|---|---|
| Language | Python | Python |
| License | MIT | MIT |
| Stars | 99 | 2.8k |
| Forks | 18 | 222 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- judgy focuses on calibrated pass-rate estimates for binary outcomes; promptbench covers model evaluation across datasets, prompting methods, and evaluation protocols.
- judgy uses human labels to estimate a judge’s true positive and true negative rates, then applies bootstrap resampling for confidence intervals; promptbench provides evaluation pipelines and analysis tools.
- judgy is suited to estimating results on larger unlabeled datasets after calibration; promptbench supports model comparisons, prompt studies, robustness testing, and multimodal evaluation.
- Both projects are Python libraries with MIT licenses, but promptbench is PyTorch-based.
- judgy is not archived, while promptbench is marked as archived; promptbench’s README also notes that its PyPI package may lag behind repository updates.
- The provided facts list 99 stars and 18 forks for judgy, compared with 2.8k stars and 222 forks for promptbench.
Choose judgy if you…
- need a corrected binary pass-rate estimate from LLM judge predictions.
- have human-labeled calibration data and want confidence bounds for an estimate on unlabeled examples.
- want a Python library focused on accounting for judge errors in pass/fail evaluation.
Choose promptbench if you…
- need to compare models across datasets, prompting methods, or evaluation protocols.
- want to examine prompt robustness, adversarial attacks, or dynamic evaluation.
- are assessing supported multimodal models on image-and-language benchmarks.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.