marker: Convert Documents into Structured Text

marker: Convert Documents into Structured Text

Summary

Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

At a glance

Language
Python
License
Apache-2.0
Stars
40.2k
Forks
2.9k
Added to OSRepos
November 9, 2025
Last analyzed
October 3, 2026
View on GitHub

Topics

Click on any tag to explore related repositories

Use at your own risk

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.

Overview

Marker is a Python document-conversion tool that turns PDFs and supported document formats into structured outputs for reading, search, or downstream processing. It addresses a common extraction problem: plain text extraction often loses page layout, tables, equations, and other document structure.

Its pipeline combines text-layer extraction with layout detection and OCR, using a vision-language model selectively. Choose it when you need a self-hosted conversion workflow and control over output formats or processing logic. For simple digital PDFs where structure is unimportant, a lighter text extractor may be sufficient.

Key Features

  • Converts PDFs and, with additional dependencies, image, PPTX, DOCX, XLSX, HTML, and EPUB files.
  • Produces Markdown, JSON, HTML, or chunked output.
  • Preserves and processes elements including tables, equations, code, links, and images.
  • Offers balanced and fast modes, with different tradeoffs in model use and conversion quality.
  • Supports batch conversion through a CLI, as well as Python and a small local API server.
  • Allows custom processors, renderers, providers, and optional LLM-based correction.

Use Cases

  • Developers preparing PDF collections for retrieval-augmented generation can export structured chunks or JSON blocks.
  • Data and research teams can extract text, equations, and tables from papers, reports, and scanned documents.
  • Organizations processing documents locally can build a conversion pipeline without sending files to a hosted service.
  • Engineers integrating document extraction into an application can use the Python API or run the provided local server for small-scale use.

Project Facts

  • Language: Python
  • License: Apache-2.0
  • Stars: 40.2k
  • Forks: 2.9k
  • Topics: none listed
  • Archived: no

Getting Started

Install the PDF package:

pip install marker-pdf

Then convert a file with marker_single /path/to/file.pdf. See the README for other formats, backend prerequisites, configuration, and usage details.

Alternatives

  • opendataloader-pdf: OpenDataLoader focuses on PDF parsing and repairing accessibility tags, while Marker converts documents into Markdown, JSON, HTML, or chunks.
  • PDF Craft: PDF Craft targets PDF-to-Markdown and EPUB conversion, especially for scanned books, rather than Marker’s broader document formats and outputs.
  • text-extract-api: text-extract-api offers an API for multiple document formats, with PII removal and LLM-backed extraction, rather than Marker’s local conversion workflow.
  • GLM-OCR: GLM-OCR provides an OCR model and SDK for text and layout extraction, while Marker packages document conversion into several structured output formats.

Considerations

  • OCR and layout processing use the Surya vision-language model through a local inference server. The README lists vLLM with Docker and the NVIDIA Container Toolkit for NVIDIA GPUs, or llama.cpp for CPU and Apple Silicon. Hardware and setup requirements depend on the selected mode and workload.
  • The project notes that complex layouts, nested tables, and forms may not convert reliably. Optional LLM processing or forced OCR can help, but may require configuring an external or local LLM service.
  • The code is Apache-2.0 licensed, but the README says model weights have a separate modified OpenRAIL-M license with limits on some commercial use. Check the model terms before deployment.
  • The included API server is described as suitable for small-scale use, not as a robust production service.

Found this useful?

Share it with someone who would like marker.

Comparisons

Source repository

Open the original repository on GitHub.

21 counted GitHub visits

View on GitHub

Related repositories

Similar repositories that may be relevant next.

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️