Open Source Document Processing Tools

Document processing covers the conversion, extraction, organization, and analysis of information in files such as PDFs, scanned pages, and office documents. These tools help make content searchable, transform files into structured formats, extract text and tables, and prepare information for indexing or further analysis. Optical character recognition can recover text from images, while parsing and conversion tools handle content and layout in digital documents.

Open source options range from focused libraries and command-line utilities to APIs and workflows that combine OCR, format conversion, and language models. When choosing a tool, consider supported formats, extraction accuracy, layout handling, language support, license, maintenance, resource requirements, and integration needs. Document processing tools are useful to developers, researchers, organizations, and anyone building searchable archives, automated workflows, or document-based applications.

10 repositories · updated August 30, 2026

pdfminer.six: Advanced PDF Parsing and Data Extraction in Python

pdfminer.six: Advanced PDF Parsing and Data Extraction in Python

pdfminer.six is a powerful, community-maintained Python library designed for extracting and analyzing text data from PDF documents. It allows users to retrieve text directly from PDF source code, including details like location, font, and color. This versatile tool also supports advanced features such as CJK languages, image extraction, and various PDF specifications.

PythonPDFParser
Added Aug 30, 2026 View details
book-to-skill: Transform Technical Books into AI Agent Skills

book-to-skill: Transform Technical Books into AI Agent Skills

book-to-skill is a Python project that converts technical books, documents, or source collections into structured agent skills. It allows AI agents like GitHub Copilot CLI or Claude Code to load content on demand, providing accurate answers without hallucination. This tool optimizes learning and reference by distilling complex information into an easily queryable format.

Agent SkillsAI AgentsLLM
Added Aug 27, 2026 View details
Article-Assistant--RAG-Telegram-Bot: AI-Powered Knowledge Base via Telegram

Article-Assistant--RAG-Telegram-Bot: AI-Powered Knowledge Base via Telegram

The Article Assistant is a sophisticated RAG (Retrieval-Augmented Generation) Telegram bot designed to create interactive knowledge bases from various documents. Users can upload PDFs or provide URLs, and the bot will provide AI-powered answers with source citations. This tool efficiently transforms static content into a dynamic, queryable resource.

PythonTelegram BotRAG
Added Apr 24, 2026 View details
RAG-Anything: The All-in-One Multimodal RAG Framework

RAG-Anything: The All-in-One Multimodal RAG Framework

RAG-Anything is a comprehensive, all-in-one Retrieval-Augmented Generation (RAG) framework designed to process and query diverse multimodal content. It seamlessly handles text, images, tables, and equations within a single integrated system, eliminating the need for multiple specialized tools. Built on LightRAG, this framework offers advanced multimodal retrieval capabilities for complex documents.

Multi Modal RAGRetrieval Augmented GenerationPython
Added Jan 31, 2026 View details
docling-api: Scalable Document to Markdown Conversion Server

docling-api: Scalable Document to Markdown Conversion Server

docling-api is a robust and scalable backend server designed for converting a wide array of document formats, including PDFs, DOCX, and images, into Markdown. Built with FastAPI, Celery, and Redis, it supports both CPU and GPU processing, making it ideal for large-scale workflows requiring efficient text, table, and image extraction, along with OCR capabilities. This service offers flexible synchronous and asynchronous API endpoints for single and batch document conversions.

PythonAPIFastAPI
Added Jan 30, 2026 View details
pdfplumber: Extracting Data from PDFs with Ease and Precision

pdfplumber: Extracting Data from PDFs with Ease and Precision

pdfplumber is a powerful Python library designed to extract detailed information from PDFs, including characters, rectangles, and lines. It excels at easily extracting text and tables, making it an invaluable tool for data analysis and automation. Built on pdfminer.six, it provides robust PDF parsing capabilities.

PDFPDF ParsingTable Extraction
Added Jan 24, 2026 View details
E2M: Convert Various File Types to Markdown for RAG and LLM Training

E2M: Convert Various File Types to Markdown for RAG and LLM Training

E2M is a Python library designed to convert diverse file types, including documents, web pages, and audio, into Markdown format. It features a robust parser-converter architecture, making it highly flexible and easy to integrate. This tool is specifically aimed at generating high-quality data for Retrieval-Augmented Generation (RAG) and large language model training.

E2mMarkdown ConversionPDF To Markdown
Added Dec 24, 2025 View details
sumy: Automatic Text Summarization for Documents and HTML Pages

sumy: Automatic Text Summarization for Documents and HTML Pages

sumy is a robust Python module designed for automatic summarization of text documents and HTML pages. It provides various summarization methods, supports multiple natural languages, and offers both a command-line utility and a flexible Python API. This versatile tool enables users to efficiently extract concise summaries from lengthy content.

PythonSummarizationNLP
Added Dec 14, 2025 View details
gptpdf: Effortlessly Parse PDFs into Markdown with GPT-4o

gptpdf: Effortlessly Parse PDFs into Markdown with GPT-4o

gptpdf is a powerful Python library that leverages large visual models like GPT-4o to accurately parse PDF documents into clean Markdown format. With just 293 lines of code, it excels at preserving typography, math formulas, tables, and images. This tool offers an efficient and cost-effective solution for converting complex PDFs.

PythonPDF ParsingMarkdown
Added Oct 24, 2025 View details
text-extract-api: Advanced Document Extraction, OCR, and PII Removal with LLMs

text-extract-api: Advanced Document Extraction, OCR, and PII Removal with LLMs

text-extract-api is a powerful API designed for extracting and parsing text from various document formats, including PDF, Word, and PPTX. It utilizes modern OCRs and Ollama-supported LLMs for highly accurate text extraction, PII removal, and conversion to structured JSON or Markdown, all while maintaining data privacy through its self-hosted architecture.

AnonymizationAPIDocument Processing
Added Oct 12, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️