gptpdf vs marker
PDF conversion with vision models compared
gptpdf and marker convert PDFs into structured text, including Markdown, while aiming to retain visual structure such as tables and equations. gptpdf relies on a compatible model API for page interpretation, whereas marker combines local extraction, OCR, layout detection, and selective vision-language processing, and supports more input and output formats.

gptpdf: Convert PDFs into Markdown with Vision Models
gptpdf turns PDF pages into Markdown using a vision-capable model, aiming to preserve layouts such as tables, formulas, and figures. It suits developers who need an API-based PDF parsing component and can provide access to a compatible model.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.
| gptpdf | marker | |
|---|---|---|
| Language | Python | Python |
| License | MIT | Apache-2.0 |
| Stars | 3.6k | 40.2k |
| Forks | 261 | 2.9k |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- gptpdf focuses on PDFs and produces Markdown plus paths to extracted images; marker also supports several other document formats and outputs JSON, HTML, and chunks.
- gptpdf sends page content for interpretation through an OpenAI-compatible model API; marker uses text extraction, layout detection, and OCR with a local inference backend.
- gptpdf requires access to a compatible multimodal model API, while marker is designed for local processing and has hardware and setup requirements that vary by mode.
- gptpdf is MIT licensed; marker's code is Apache-2.0 licensed, and its model weights have separate modified OpenRAIL-M terms.
- The supplied project figures list gptpdf at 3.6k stars and 261 forks, and marker at 40.2k stars and 2.9k forks.
Choose gptpdf if you…
- need a PDF-to-Markdown component that uses a compatible model API to interpret page content.
- want to configure the model API, prompts, and worker threads for PDF parsing.
- can accept per-page model costs and review output where exact transcription or layout fidelity matters.
Choose marker if you…
- need conversion of PDFs and other supported formats into Markdown, JSON, HTML, or chunks.
- want a locally run workflow using text extraction, OCR, layout detection, and selective vision-language processing.
- need CLI batch conversion, a Python API, or custom processors and renderers.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.