Repository History
2 repositories tagged with Text Extraction
Topic: Text Extraction

pdf-inspector: Fast Rust Library for PDF Classification and Text Extraction
pdf-inspector is a high-performance Rust library designed for intelligent PDF processing. It excels at classifying PDFs as text-based or scanned, extracting text with position awareness, and converting content to clean Markdown. This library enables smart routing decisions, significantly reducing the need for expensive OCR services for many documents.
Analyzed Aug 26, 2026
View Details

Trafilatura: Advanced Web Scraping and Text Extraction in Python
Trafilatura is a robust Python package and command-line tool designed for gathering text and metadata from the web. It simplifies web crawling, scraping, and content extraction, transforming raw HTML into structured data. Widely adopted by major companies and institutions, it offers high efficiency and accuracy for various text processing needs.
Analyzed May 1, 2026
View Details
Previous Page 1 Next