datatrove: Build Large-Scale Text Data Pipelines

Summary
DataTrove is a Python library for processing, filtering, and deduplicating text datasets at scale. It combines reusable pipeline blocks with local and cluster executors, making it useful for data preparation workflows such as building LLM training corpora.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 3.4k
- Forks
- 310
- Added to OSRepos
- January 27, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
DataTrove is a pipeline framework for turning large collections of text files into cleaned, filtered, deduplicated datasets. It is designed to replace one-off processing scripts with reusable blocks that can be composed into a pipeline and run across different execution environments.
The project is aimed at teams working with large text corpora, especially data engineers and machine-learning practitioners preparing training data. Its main value is the combination of data-processing components, distributed task execution, and resumable jobs, rather than a single turnkey dataset-cleaning recipe.
Key Features
- Compose reader, extractor, filter, statistics, tokenization, deduplication, and writer blocks into pipelines.
- Run the same pipeline locally or through Local, Slurm, or Ray executors.
- Split work into tasks and resume by skipping tasks marked as completed after a relaunch.
- Read and write across local, remote, and cloud-backed filesystems through fsspec.
- Process common text-data formats and extract text from raw sources such as webpage HTML.
- Collect dataset statistics across shards and merge results.
- Add custom processing functions or blocks that operate on DataTrove documents.
- Use optional inference components for synthetic data generation with supported inference servers or compatible endpoints.
Use Cases
- Data teams preparing web-scale text for language-model training can combine extraction, quality filters, deduplication, and output steps in one pipeline.
- Researchers processing Common Crawl or other large file collections can distribute files across Slurm tasks and retry incomplete work.
- Dataset maintainers can profile text characteristics, such as length and language scores, before publishing a cleaned dataset.
- Machine-learning teams can generate synthetic text using the optional inference pipeline, then write results to intermediate storage for further processing.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 3.4k
- Forks: 310
- Archived: false
Getting Started
Python 3.10 or newer is required. Install the base project and choose optional dependencies for the formats and execution backends you need:
uv sync
For example, add extras with uv sync --extra processing --extra s3. See the README for installation options, examples, and executor configuration.
Considerations
- DataTrove is a framework, not an end-to-end data-cleaning policy. You need to choose and configure the pipeline steps appropriate for your corpus.
- Effective parallelism depends on having enough input files to distribute across tasks. The README notes that a single file is not automatically split into multiple parts.
- Slurm, Ray, S3, inference, and other capabilities require their corresponding optional dependencies and environment setup.
- The Hugging Face Jobs executor is explicitly experimental and may change or be removed without notice.
- The project requires Python 3.10 or newer.
Source repository
Open the original repository on GitHub.
17 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design
October 3, 2026
The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner
October 2, 2026
OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

Shepherd: Reversible Execution Traces for Programmable Meta-Agents
October 2, 2026
Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
October 1, 2026
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.