Document Parsing
Document parsing turns files such as PDFs, office documents, and scanned pages into structured, searchable data. It identifies text, tables, images, layout, and metadata so information can be indexed, analyzed, or passed to downstream applications. Parsing helps address the difficulty of working with varied file formats and extracting reliable content from documents that were not designed for automated processing.
Open source tools in this area range from format converters and layout analyzers to optical character recognition systems and pipelines that prepare content for language models. When choosing a tool, consider the file types and languages it supports, extraction accuracy, accessibility features, output formats, and hardware requirements. Review its license, documentation, integration options, and maintenance activity. These tools are useful to developers, researchers, and organizations building search, data extraction, or document automation workflows.
2 repositories · updated October 3, 2026

docling: Convert Documents into Structured Content
Docling converts PDFs and many other document formats into structured representations and exports such as Markdown and JSON. It is suited to developers building document ingestion workflows for search, analytics, and generative AI, including local processing of sensitive files.

docling-api: Convert Documents to Markdown Through an API
docling-api is a self-hostable FastAPI service that converts documents and images into Markdown using Docling. It suits teams that need synchronous or queued batch processing, with CPU and GPU deployment options.