text-extract-api vs marker
Self-hosted document extraction tools compared
text-extract-api and marker convert documents into structured outputs for downstream use. text-extract-api centers on an API with OCR and optional Ollama processing, while marker focuses on document layout conversion and offers several output formats.

text-extract-api: Extract Text and Data from Documents
A self-hostable API that turns PDFs, images, and Office files into Markdown or structured JSON using OCR and Ollama models. It suits teams that need document processing and PII removal with control over where files are processed.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.
| text-extract-api | marker | |
|---|---|---|
| Language | Python | Python |
| License | MIT | Apache-2.0 |
| Stars | 3.2k | 40.2k |
| Forks | 279 | 2.9k |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- text-extract-api accepts PDFs, images, and Office documents through an HTTP API or CLI; marker also supports formats such as HTML and EPUB with additional dependencies.
- text-extract-api outputs Markdown or structured JSON; marker can also produce HTML and chunks.
- text-extract-api combines OCR strategies with optional Ollama processing; marker uses text-layer extraction, layout detection, OCR, and selective vision-language model processing.
- text-extract-api includes asynchronous tasks, Redis caching, and configurable result storage; marker offers batch conversion, custom processors, and renderers.
- text-extract-api is MIT licensed; marker is Apache-2.0 licensed, and its model weights have separate modified OpenRAIL-M terms.
- text-extract-api documents an API service stack with FastAPI, Celery, Redis, and Ollama; marker's included API server is described as suitable for small-scale use.
Choose text-extract-api if you…
- need an HTTP API or CLI for extracting documents into Markdown or JSON.
- want asynchronous processing, OCR-result caching, or configurable storage profiles.
- need optional Ollama processing, including prompt-based PII removal.
Choose marker if you…
- need output as HTML or chunks as well as Markdown or JSON.
- want to process document structure such as tables, equations, code, links, and images.
- can configure its local inference backend and want options for batch conversion or custom processing logic.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.