Open Source NLP Projects

Discover 24 open source NLP repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. NLP projects here are most often combined with Python, Library and Machine Learning. Last updated October 3, 2026.

24 repositories · updated October 3, 2026

unstructured: Turn Documents Into Structured Data

unstructured: Turn Documents Into Structured Data

Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

PythonLLMData Extraction
Added Feb 10, 2026 View details
LLMSanitize: Detect Contamination in NLP Data and LLMs

LLMSanitize: Detect Contamination in NLP Data and LLMs

LLMSanitize brings together methods for checking whether NLP datasets or language models may be contaminated by training data. It is aimed at researchers and evaluators who need to assess benchmark reliability across open- and closed-data settings.

PythonLibraryMachine Learning
Added Feb 9, 2026 View details
datatrove: Build Large-Scale Text Data Pipelines

datatrove: Build Large-Scale Text Data Pipelines

DataTrove is a Python library for processing, filtering, and deduplicating text datasets at scale. It combines reusable pipeline blocks with local and cluster executors, making it useful for data preparation workflows such as building LLM training corpora.

PythonLibraryData Science
Added Jan 27, 2026 View details
sumy: Summarize Text and HTML Documents

sumy: Summarize Text and HTML Documents

sumy is a Python library and command-line tool that extracts concise summaries from plain text and HTML. It offers several extractive summarization methods, language tokenization support, and tools for evaluating summaries.

PythonNLPNatural Language Processing
Added Dec 14, 2025 View details
Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression

Toolkit-for-Prompt-Compression: Evaluate and Apply Prompt Compression

PCToolkit is a Python toolkit for applying and evaluating prompt-compression methods for large language models. It brings five compressors, datasets, and evaluation metrics behind modular interfaces, making it useful for comparing methods across language tasks.

PythonAILLM
Added Dec 13, 2025 View details
whisper-web: Transcribe Speech in Your Browser

whisper-web: Transcribe Speech in Your Browser

Whisper Web uses Transformers.js to run speech recognition in a browser. It suits people who want to try transcription without setting up a separate server, though browser support and performance may vary.

TypeScriptJavaScriptMachine Learning
Added Dec 5, 2025 View details
LlamaFactory: Fine-Tune Large Language and Vision Models

LlamaFactory: Fine-Tune Large Language and Vision Models

LlamaFactory provides CLI and web interfaces for fine-tuning a broad range of language and vision models. It supports parameter-efficient methods and preference training, with workflows for training, inference, and model export.

PythonLLMMachine Learning
Added Nov 8, 2025 View details
text-generation-inference: Serve Large Language Models

text-generation-inference: Serve Large Language Models

A toolkit for serving large language models through text-generation APIs. It provides GPU-oriented inference features such as continuous batching, token streaming, tensor parallelism, and quantization, but is now in maintenance mode.

PythonLLMMachine Learning
Added Nov 4, 2025 View details
textdistance: Compare Sequences with Distance Algorithms

textdistance: Compare Sequences with Distance Algorithms

TextDistance is a Python library for comparing strings and other sequences with a common interface to more than 30 distance and similarity algorithms. Use its pure-Python implementations or optional external libraries when speed matters.

PythonLibraryText Processing
Added Oct 31, 2025 View details
python-ftfy: Repair Mojibake and Unicode Text

python-ftfy: Repair Mojibake and Unicode Text

ftfy is a Python library that repairs mojibake and other common Unicode glitches in text. It suits developers cleaning text from mixed or unreliable sources, with conservative fixes designed to avoid changing text that is already valid.

PythonLibraryText Processing
Added Oct 21, 2025 View details
newspaper: Extract News Articles and Metadata in Python

newspaper: Extract News Articles and Metadata in Python

Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.

PythonLibraryWeb Scraping
Added Oct 13, 2025 View details
FinGPT: Adapt Language Models for Financial Tasks

FinGPT: Adapt Language Models for Financial Tasks

FinGPT provides financial datasets, fine-tuned language models, benchmarks, and workflows for tasks such as sentiment analysis and forecasting. It suits researchers and developers adapting open models to finance, with local GPU inference or supported cloud APIs.

AILarge Language ModelsMachine Learning
Added Oct 12, 2025 View details
Previous Page 2 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️