marker vs xberg
Document extraction projects compared
marker and xberg extract structured content from documents for downstream workflows such as search and RAG. marker focuses on document conversion with local layout processing, while xberg offers a broader extraction engine with multiple interfaces and language bindings.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

xberg: Extract Text and Structure from Documents
Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.
| marker | xberg | |
|---|---|---|
| Language | Python | Rust |
| License | Apache-2.0 | MIT |
| Stars | 40.2k | 9.4k |
| Forks | 2.9k | 588 |
| Last analyzed | Oct 3, 2026 | Oct 3, 2026 |
Key differences
- marker is a Python tool centered on converting PDFs and supported document formats into Markdown, JSON, HTML, or chunks; xberg is a Rust engine for documents, archives, audio, web content, and source code.
- marker uses text extraction, layout detection, OCR, and selective vision-language model processing; xberg provides several OCR backends and optional features for capabilities such as transcription and URL ingestion.
- marker provides a CLI, Python API, and small local API server; xberg provides a library, CLI, REST API, MCP server, and bindings for 15 languages.
- marker is licensed under Apache-2.0, while xberg is licensed under MIT. marker also notes separate modified OpenRAIL-M terms for its model weights.
- marker’s documentation cautions that complex layouts may not convert reliably and describes its API server as suitable for small-scale use; xberg’s integration details vary by binding and some capabilities require optional features.
- marker reports 40.2k stars and 2.9k forks; xberg reports 9.4k stars and 588 forks.
Choose marker if you…
- need structured conversion of PDFs and supported document formats into Markdown, JSON, HTML, or chunks.
- want to build a self-hosted Python document-processing workflow with custom processors or renderers.
- process papers or scanned documents and need to preserve elements such as tables, equations, and images.
Choose xberg if you…
- need extraction across varied content, including archives, audio, web content, or source code.
- want to integrate document extraction from a language other than Rust using available bindings.
- need a REST API or MCP server, or want options such as syntax-aware code chunking and transcription when enabled.
This comparison is generated with AI from the OSRepos analyses of both projects. Always check each project's repository and documentation before choosing.