Our take
- Activity
- 75
- Community
- 86
- Issues
- 72
- Pull requests
- 73
- More contributors than 77% of the projects we track
Overview
MarkItDown is a Python tool for extracting document content and structure into Markdown. It is aimed at developers building LLM or text-analysis pipelines, where readable headings, lists, tables, and links are more useful than a high-fidelity visual reproduction of the original file.
It supports a range of document and media inputs through built-in converters and optional integrations. Use it when downstream systems need text and basic structure; choose a layout-focused conversion service or another tool when visual fidelity is the priority.
Key Features
- Converts formats including PDF, Word, PowerPoint, Excel, HTML, images, audio, EPUB, ZIP, and YouTube URLs.
- Preserves useful document structure such as headings, lists, tables, and links in Markdown output.
- Provides both a Python API and a command-line interface, including support for piping input and output.
- Offers optional dependencies that let users install support for selected formats rather than every format.
- Supports third-party plugins, which are disabled by default and can be enabled explicitly.
- Integrates with Azure Document Intelligence and Azure Content Understanding for cloud-based extraction.
- Can use an LLM client for image descriptions in supported image and PowerPoint conversions.
Use Cases
- LLM application developers can convert office documents into text that is easier to include in retrieval or analysis pipelines.
- Data teams can extract text and tabular content from mixed-format files for downstream processing.
- Developers can add document conversion to a Python application or automate one-off conversions from the terminal.
- Teams processing scanned or complex documents can evaluate the Azure integrations when local extraction is insufficient.
Getting Started
Install the package with all optional dependencies, then convert a file:
pip install 'markitdown[all]'
markitdown path-to-file.pdf -o document.md
For selected format dependencies, Python API examples, and integration details, see the README.
Alternatives
- docling: Docling handles a broad range of document formats but emphasizes structured representations and exports, while MarkItDown focuses on Markdown conversion.
- pdf-craft: PDF Craft specializes in OCR-based conversion of scanned documents to Markdown or EPUB, while MarkItDown supports a wider range of file formats.
- gptpdf: gptpdf uses a vision-capable model to convert PDF pages to Markdown and preserve layout, while MarkItDown converts many formats with optional dependencies.
- pdf-inspector: pdf-inspector focuses on PDF classification and positioned text extraction, routing scanned pages to OCR; MarkItDown targets Markdown conversion across formats.
| Project | Language | License | Stars | Status |
|---|---|---|---|---|
| markitdown | Python | MIT | 188k | Not checked yet |
| docling | Python | MIT | 68.3k | Active |
| pdf-craft | Python | MIT | 6.3k | Active |
| gptpdf | Python | MIT | 3.6k | Inactive |
| pdf-inspector | Rust | MIT | 19.5k | Active |
Considerations
- The output is designed for text analysis, not high-fidelity visual document reproduction. Check the result against the source when layout or appearance matters.
- Python 3.10 through 3.14 is required. Format support may require optional dependencies.
- Azure-based conversion requires an endpoint and makes billable API calls. The project also notes that Content Understanding is needed for video conversion.
- Conversion performs I/O with the privileges of the calling process. Validate untrusted inputs and restrict file paths, URI schemes, and network access in hosted or server-side applications.
- The README describes image descriptions through an LLM client for supported formats. Treat this as an optional integration, not a requirement for basic conversion.
Found this useful?
Share it with someone who would like markitdown.