marker vs pdf-inspector
PDF and document extraction tools compared
marker converts PDFs and other supported documents into structured formats, including Markdown, JSON, HTML, and chunks. pdf-inspector focuses on classifying PDFs, extracting positioned text, and routing pages that need OCR; the projects also differ in language and deployment approach.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

pdf-inspector: Classify PDFs and Extract Text as Markdown
pdf-inspector is a Rust library and set of bindings for classifying PDFs, extracting positioned text, and converting documents to Markdown. It helps processing pipelines route scanned pages to OCR while handling text-based pages locally.
| marker | pdf-inspector | |
|---|---|---|
| Language | Python | Rust |
| License | Apache-2.0 | MIT |
| Stars | 40.2k | 19.5k |
| Forks | 2.9k | 1.3k |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- marker is a Python tool for PDFs and additional document formats, while pdf-inspector is a Rust PDF library with Python, Node.js, browser WebAssembly, and CLI interfaces.
- marker uses layout detection, OCR, and selective vision-language model processing; pdf-inspector classifies pages and extracts text locally, with OCR as an optional path in supported native integrations.
- marker outputs Markdown, JSON, HTML, or chunks; pdf-inspector focuses on structured Markdown and positioned text, including font and coordinate information.
- marker uses the Apache-2.0 code license, with separate model-weight terms; pdf-inspector uses the MIT license.
- marker supports local inference through a model server and offers a small-scale API server; pdf-inspector's lightweight default Rust and browser builds do not include OCR.
- marker offers balanced and fast conversion modes and customizable processors; pdf-inspector emphasizes page classification, confidence, and selective OCR routing.
Choose marker if you…
- need to convert supported documents beyond PDFs into several structured output formats.
- want to preserve elements such as equations and images, or export chunks for downstream workflows.
- can run local models and need customizable conversion processing.
Choose pdf-inspector if you…
- need to classify PDFs and route only pages requiring OCR.
- want positioned text, reading order, and Markdown from a Rust-based library or its language bindings.
- need local extraction for text-based PDFs with optional OCR in supported native integrations.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.