Open Source OCR Projects

Discover 18 open source OCR repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. OCR projects here are most often combined with Data Extraction, Python and PDF. Last updated October 3, 2026.

18 repositories · updated October 3, 2026

docling-api: Convert Documents to Markdown Through an API

docling-api: Convert Documents to Markdown Through an API

docling-api is a self-hostable FastAPI service that converts documents and images into Markdown using Docling. It suits teams that need synchronous or queued batch processing, with CPU and GPU deployment options.

PythonAPIDocument Parsing
Added Jan 30, 2026 View details
papermerge: Organize and Search Scanned Documents

papermerge: Organize and Search Scanned Documents

Papermerge is a web-based document management system for organizing scanned archives. It uses OCR and full-text search to make documents easier to find, but this repository is archived and development has moved to papermerge-core.

PythonDjangoOCR
Added Jan 8, 2026 View details
xberg: Extract Text and Structure from Documents

xberg: Extract Text and Structure from Documents

Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

RustData ExtractionLibrary
Added Dec 30, 2025 View details
marker: Convert Documents into Structured Text

marker: Convert Documents into Structured Text

Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

PythonAIPDF
Added Nov 9, 2025 View details
Ollama-OCR: Extract Text from Images and PDFs with Vision Models

Ollama-OCR: Extract Text from Images and PDFs with Vision Models

Ollama-OCR uses vision-language models served by Ollama to extract text and structured content from images and PDFs. It offers a Python package for single or batch processing and a Streamlit interface for interactive use.

PythonAIComputer Vision
Added Oct 12, 2025 View details
text-extract-api: Extract Text and Data from Documents

text-extract-api: Extract Text and Data from Documents

A self-hostable API that turns PDFs, images, and Office files into Markdown or structured JSON using OCR and Ollama models. It suits teams that need document processing and PII removal with control over where files are processed.

PythonAPIOCR
Added Oct 12, 2025 View details
Previous Page 2 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️