LLMSanitize: Detect Contamination in NLP Data and LLMs

LLMSanitize: Detect Contamination in NLP Data and LLMs

Summary

LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.

At a glance

Language
Python
License
Apache-2.0
Stars
61
Forks
6
Added to OSRepos
February 9, 2026
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

LLMSanitize is a Python library for investigating contamination between evaluation data and the data used to train or refine language models. It packages a range of published detection approaches, helping NLP researchers and model evaluators compare methods rather than build each one from scratch.

The methods cover both open-data checks, such as string matching and embedding similarity, and closed-data checks that query or inspect a model. It is most useful when evaluating whether benchmark results could be affected by exposure to test examples, and when the evaluator can provide the required model access and infrastructure.

Key Features

  • Supports open-data checks using string matching and embedding similarity.
  • Includes closed-data approaches based on likelihood, memorization, prompting, and model completion.
  • Covers black-box and white-box model access patterns, depending on the method.
  • Provides test scripts for running detection methods on datasets such as Hellaswag.
  • Can be used through shell scripts or called from Python code.
  • Uses a vLLM instance for several methods that require model inference.

Use Cases

  • NLP researchers can investigate possible overlap between benchmark examples and model training data before interpreting evaluation results.
  • Benchmark maintainers can compare contamination-checking approaches for datasets they publish or curate.
  • Model evaluators can assess closed-data contamination when they have access to a model and the method's required inference setup.
  • Research teams studying data memorization can run multiple published detection methods within one library.

Project Facts

  • Language: Python
  • License: Apache-2.0
  • Stars: 61
  • Forks: 6
  • Archived: No

Getting Started

The README specifies Python 3.9 and CUDA 11.8 as the tested setup. Install the package with:

pip install llmsanitize

See the repository README for environment setup, method-specific scripts, and usage details.

Considerations

  • The documented setup targets Python 3.9 and CUDA 11.8, and the project notes that it uses vLLM 0.3.3.
  • Several methods require launching a vLLM inference server and specifying a port and model name. Other methods may require white-box access to model likelihoods or a model path.
  • The available approaches have different data and access assumptions, so a method that fits one evaluation setting may not fit another.
  • The repository reports one open issue and was last pushed on 2024-08-13, based on the provided repository metadata.

Source repository

Open the original repository on GitHub.

20 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️