Markdown Conversion
Markdown conversion transforms content from formats such as PDFs, word-processing files, web pages, and images into Markdown, a lightweight format for structured text. It helps make information easier to edit, publish, search, and use in documentation or AI workflows. Conversion can involve more than extracting words: preserving headings, tables, links, reading order, and other document structure is often important, especially for complex or scanned files.
Open source tools in this area range from format-specific parsers and conversion libraries to APIs and services for processing documents at scale. When choosing one, consider supported formats, output quality, maintenance activity, license, runtime requirements, and how well it integrates with your existing workflows. These tools can be useful to developers building content pipelines, teams migrating documents, and anyone who needs readable, reusable text from varied sources.
3 repositories · updated August 26, 2026

pdf-inspector: Fast Rust Library for PDF Classification and Text Extraction
pdf-inspector is a high-performance Rust library designed for intelligent PDF processing. It excels at classifying PDFs as text-based or scanned, extracting text with position awareness, and converting content to clean Markdown. This library enables smart routing decisions, significantly reducing the need for expensive OCR services for many documents.

Docling: Streamlining Document Processing for Generative AI
Docling is a powerful Python library designed to simplify document processing and prepare diverse formats for generative AI applications. It excels at parsing various document types, including advanced PDF understanding, and offers seamless integrations with popular AI frameworks. With Docling, developers can efficiently extract, transform, and utilize document content for their AI models.

E2M: Convert Various File Types to Markdown for RAG and LLM Training
E2M is a Python library designed to convert diverse file types, including documents, web pages, and audio, into Markdown format. It features a robust parser-converter architecture, making it highly flexible and easy to integrate. This tool is specifically aimed at generating high-quality data for Retrieval-Augmented Generation (RAG) and large language model training.