Open Source OCR Projects
Optical character recognition (OCR) converts text in images and scanned documents into machine-readable text. It helps make printed or handwritten content searchable, editable, and accessible to software, reducing manual transcription and making information easier to process. Modern OCR systems may also detect page layout, tables, and other document structure, which can be important when preserving meaning and formatting matters.
Open source OCR tools range from text-recognition libraries and command-line utilities to document-processing pipelines and systems that use machine learning. When choosing one, consider language and handwriting support, accuracy on your document types, output formats, processing speed, hardware requirements, licensing, maintenance, and integration with existing workflows. These tools are useful to developers, researchers, organizations, and archivists working with scanned records, PDFs, forms, or image collections.
15 repositories · updated October 3, 2026

AWESOME-OCR-LLM: Curated Reading List for OCR in the LLM Era
AWESOME-OCR-LLM is a continuously updated reading list focusing on Optical Character Recognition (OCR) in the era of large language models (LLMs). It covers key areas like document parsing, understanding, visual text generation, and benchmarks, highlighting research from the past five years. This resource is invaluable for anyone tracking the rapid advancements in multimodal document AI.

Qwen3-VL: Understand Images, Video, and Text with Multimodal Models
Qwen3-VL is a family of multimodal language models for interpreting images, video, and text. The repository provides inference examples, deployment guidance, and cookbooks for tasks such as OCR, spatial reasoning, and visual agents.

opendataloader-pdf: Extract Structured Data and Accessibility Tags from PDFs
OpenDataLoader PDF parses digital, scanned, and tagged PDFs into structured formats for AI and document workflows. It also automates conversion of untagged PDFs into Tagged PDFs, with optional hybrid processing for complex documents.

GLM-OCR: Recognize Text and Structure in Documents
GLM-OCR is a multimodal OCR model and SDK for extracting text and layout from complex documents. Use it through a hosted API or deploy the pipeline with supported inference servers for local control.

OmniParse: Ingest, Parse, and Optimize Data for GenAI Frameworks
OmniParse is a powerful platform designed to ingest, parse, and optimize any unstructured data, from documents to multimedia, into structured, actionable formats. It enhances compatibility with GenAI frameworks, preparing data for applications like RAG and fine-tuning. This tool simplifies the complex process of data preparation for AI, making it accessible and efficient.

ImageToolbox: Edit, Convert, and Process Images on Android
ImageToolbox is a feature-rich Android app for photo editing, image conversion, OCR, PDF tasks, and other media workflows. It suits users who want many image utilities in one place, with distribution through Google Play, F-Droid, and GitHub releases.

PaddleOCR: A Powerful OCR Toolkit for Structured Document Data
PaddleOCR is an industry-leading, production-ready OCR and document AI engine that transforms any PDF or image document into structured, AI-friendly data. It offers end-to-end solutions from text extraction to intelligent document understanding, supporting over 100 languages with high accuracy and efficiency.

PDF Craft: Convert Scanned PDF Books to Markdown and EPUB
PDF Craft is a Python library designed to convert PDF files, especially scanned books, into various formats like Markdown and EPUB. Leveraging DeepSeek OCR, it accurately extracts text, tables, and formulas while preserving document structure. The project offers a fast, local conversion process, making it ideal for digitizing complex documents.

Unstructured: Open-Source Pre-Processing for Complex Document Data
The `unstructured` library is an open-source ETL solution designed to convert complex, unstructured documents into clean, structured data. It streamlines the data processing workflow for language models, offering tools for ingesting and pre-processing various document types like PDFs, HTML, and Word documents. This library simplifies the transformation of raw information into formats suitable for advanced AI applications.

docling-api: Scalable Document to Markdown Conversion Server
docling-api is a robust and scalable backend server designed for converting a wide array of document formats, including PDFs, DOCX, and images, into Markdown. Built with FastAPI, Celery, and Redis, it supports both CPU and GPU processing, making it ideal for large-scale workflows requiring efficient text, table, and image extraction, along with OCR capabilities. This service offers flexible synchronous and asynchronous API endpoints for single and batch document conversions.

Papermerge DMS: Open Source Document Management for Digital Archives
Papermerge DMS is an open-source document management system specifically designed for scanned documents and digital archives. It leverages OCR technology to extract and index text, enabling full-text search and efficient organization. With a modern web UI, it provides a desktop-like experience for managing various document formats.

Kreuzberg: A Polyglot Document Intelligence Framework with a Rust Core
Kreuzberg is a powerful polyglot document intelligence framework built with a high-performance Rust core. It enables extraction of text, metadata, and structured information from over 50 file formats, including PDFs, Office documents, and images. Developers can leverage Kreuzberg across multiple languages like Rust, Python, Ruby, Go, and Node.js, or utilize it via CLI, REST API, or MCP server.