Open Source Text Extraction Tools

Text extraction is the process of turning content in documents, web pages, and other files into usable text and structured data. It helps make information searchable, supports analysis and summarization, and reduces the effort needed to reuse content across formats. Some methods preserve layout or metadata, while others focus on collecting clean text from complex sources such as scanned documents or web pages.

Open source tools in this area include parsers, web content extractors, optical character recognition systems, and format converters. When choosing one, consider supported input types, extraction accuracy, language coverage, output structure, dependencies, and how actively it is maintained. Check the license and how well it integrates with your existing workflows. These tools are useful to developers, researchers, publishers, and teams that need to process or analyze large collections of content.

4 repositories · updated October 3, 2026

pdf-inspector: Fast Rust Library for PDF Classification and Text Extraction

pdf-inspector: Fast Rust Library for PDF Classification and Text Extraction

pdf-inspector is a high-performance Rust library designed for intelligent PDF processing. It excels at classifying PDFs as text-based or scanned, extracting text with position awareness, and converting content to clean Markdown. This library enables smart routing decisions, significantly reducing the need for expensive OCR services for many documents.

RustPDFText Extraction
Added Aug 26, 2026 View details
Trafilatura: Advanced Web Scraping and Text Extraction in Python

Trafilatura: Advanced Web Scraping and Text Extraction in Python

Trafilatura is a robust Python package and command-line tool designed for gathering text and metadata from the web. It simplifies web crawling, scraping, and content extraction, transforming raw HTML into structured data. Widely adopted by major companies and institutions, it offers high efficiency and accuracy for various text processing needs.

PythonWeb ScrapingText Extraction
Added May 1, 2026 View details
E2M: Convert Various File Types to Markdown for RAG and LLM Training

E2M: Convert Various File Types to Markdown for RAG and LLM Training

E2M is a Python library designed to convert diverse file types, including documents, web pages, and audio, into Markdown format. It features a robust parser-converter architecture, making it highly flexible and easy to integrate. This tool is specifically aimed at generating high-quality data for Retrieval-Augmented Generation (RAG) and large language model training.

E2mMarkdown ConversionPDF To Markdown
Added Dec 24, 2025 View details
sumy: Summarize Text and HTML Documents

sumy: Summarize Text and HTML Documents

sumy is a Python library and command-line tool that extracts concise summaries from plain text and HTML. It offers several extractive summarization methods, language tokenization support, and tools for evaluating summaries.

PythonNLPNatural Language Processing
Added Dec 14, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️