Open Source NLP Projects
Discover 24 open source NLP repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. NLP projects here are most often combined with Python, Library and Machine Learning. Last updated October 3, 2026.
24 repositories · updated October 3, 2026

unstructured: Turn Documents Into Structured Data
Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

LLMSanitize: Detect Contamination in NLP Data and LLMs
LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.

datatrove: Build Large-Scale Text Data Pipelines
DataTrove is a Python library for processing, filtering, and deduplicating text datasets at scale. It combines reusable pipeline blocks with local and cluster executors, making it useful for data preparation workflows such as building LLM training corpora.

sumy: Summarize Text and HTML Documents
sumy is a Python library and command-line tool that extracts concise summaries from plain text and HTML. It offers several extractive summarization methods, language tokenization support, and tools for evaluating summaries.

Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression
PCToolkit is a Python toolkit for applying and evaluating prompt-compression methods for large language models. It brings five compressors, datasets, and evaluation metrics behind modular interfaces, making it useful for comparing methods across language tasks.

whisper-web: Transcribe Speech in Your Browser
Whisper Web uses Transformers.js to run speech recognition in a browser. It suits people who want to try transcription without setting up a separate server, though browser support and performance may vary.

LlamaFactory: Fine-Tune Large Language and Vision Models
LlamaFactory provides CLI and web interfaces for fine-tuning a broad range of language and vision models. It supports parameter-efficient methods and preference training, with workflows for training, inference, and model export.

text-generation-inference: Serve Large Language Models
A toolkit for serving large language models through text-generation APIs. It provides GPU-oriented inference features such as continuous batching, token streaming, tensor parallelism, and quantization, but is now in maintenance mode.

textdistance: Compare Sequences with Distance Algorithms
TextDistance is a Python library for comparing strings and other sequences with a common interface to more than 30 distance and similarity algorithms. Use its pure-Python implementations or optional external libraries when speed matters.

python-ftfy: Repair Mojibake and Unicode Text
ftfy is a Python library that repairs mojibake and other common Unicode glitches in text. It suits developers cleaning text from mixed or unreliable sources, with conservative fixes designed to avoid changing text that is already valid.

newspaper: Extract News Articles and Metadata in Python
Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.

FinGPT: Adapt Language Models for Financial Tasks
FinGPT provides financial datasets, fine-tuned language models, benchmarks, and workflows for tasks such as sentiment analysis and forecasting. It suits researchers and developers adapting open models to finance, with local GPU inference or supported cloud APIs.