marker: Convert Documents into Structured Text

Summary
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.
At a glance
- Language
- Python
- License
- Apache-2.0
- Stars
- 40.2k
- Forks
- 2.9k
- Added to OSRepos
- November 9, 2025
- Last analyzed
- October 3, 2026
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Marker is a Python document-conversion tool that turns PDFs and supported document formats into structured outputs for reading, search, or downstream processing. It addresses a common extraction problem: plain text extraction often loses page layout, tables, equations, and other document structure.
Its pipeline combines text-layer extraction with layout detection and OCR, using a vision-language model selectively. Choose it when you need a self-hosted conversion workflow and control over output formats or processing logic. For simple digital PDFs where structure is unimportant, a lighter text extractor may be sufficient.
Key Features
- Converts PDFs and, with additional dependencies, image, PPTX, DOCX, XLSX, HTML, and EPUB files.
- Produces Markdown, JSON, HTML, or chunked output.
- Preserves and processes elements including tables, equations, code, links, and images.
- Offers balanced and fast modes, with different tradeoffs in model use and conversion quality.
- Supports batch conversion through a CLI, as well as Python and a small local API server.
- Allows custom processors, renderers, providers, and optional LLM-based correction.
Use Cases
- Developers preparing PDF collections for retrieval-augmented generation can export structured chunks or JSON blocks.
- Data and research teams can extract text, equations, and tables from papers, reports, and scanned documents.
- Organizations processing documents locally can build a conversion pipeline without sending files to a hosted service.
- Engineers integrating document extraction into an application can use the Python API or run the provided local server for small-scale use.
Project Facts
- Language: Python
- License: Apache-2.0
- Stars: 40.2k
- Forks: 2.9k
- Topics: none listed
- Archived: no
Getting Started
Install the PDF package:
pip install marker-pdf
Then convert a file with marker_single /path/to/file.pdf. See the README for other formats, backend prerequisites, configuration, and usage details.
Alternatives
- opendataloader-pdf: OpenDataLoader focuses on PDF parsing and repairing accessibility tags, while Marker converts documents into Markdown, JSON, HTML, or chunks.
- PDF Craft: PDF Craft targets PDF-to-Markdown and EPUB conversion, especially for scanned books, rather than Marker’s broader document formats and outputs.
- text-extract-api: text-extract-api offers an API for multiple document formats, with PII removal and LLM-backed extraction, rather than Marker’s local conversion workflow.
- GLM-OCR: GLM-OCR provides an OCR model and SDK for text and layout extraction, while Marker packages document conversion into several structured output formats.
Considerations
- OCR and layout processing use the Surya vision-language model through a local inference server. The README lists vLLM with Docker and the NVIDIA Container Toolkit for NVIDIA GPUs, or llama.cpp for CPU and Apple Silicon. Hardware and setup requirements depend on the selected mode and workload.
- The project notes that complex layouts, nested tables, and forms may not convert reliably. Optional LLM processing or forced OCR can help, but may require configuring an external or local LLM service.
- The code is Apache-2.0 licensed, but the README says model weights have a separate modified OpenRAIL-M license with limits on some commercial use. Check the model terms before deployment.
- The included API server is described as suitable for small-scale use, not as a robust production service.
Found this useful?
Share it with someone who would like marker.
Comparisons
Source repository
Open the original repository on GitHub.
21 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

agentevals: Evaluate AI Agents from OpenTelemetry Traces
October 4, 2026
agentevals scores AI agent behavior from existing OpenTelemetry traces, without rerunning agents or making extra model calls. It suits teams building instrumented agents that need local evaluation, golden-set checks, or CI quality gates.

web-design: Create Consistent Web Pages with a Claude Code Skill
October 3, 2026
web-design is a Claude Code skill that turns product briefs, reference URLs, or screenshots into an editable design specification before generating web code. It is suited to developers and designers who want a repeatable, spec-led workflow for building consistent pages.

oomwoo: Build a DIY Robot Vacuum
October 2, 2026
OOMWOO is a planned, hackable robot vacuum built around Raspberry Pi, ROS2 and 2D LiDAR. It is aimed at makers who want to build and customize a locally controlled vacuum, but its hardware and build instructions are still in development.

shepherd: Supervise Agents with Reversible Execution Traces
October 2, 2026
Shepherd records agent work as inspectable, reversible execution traces and keeps changes as proposals for review. It is aimed at developers building systems that supervise, replay, or manage the work of other agents.