docling: Convert Documents into Structured Content

Summary
Docling converts PDFs and many other document formats into structured representations and exports such as Markdown and JSON. It is suited to developers building document ingestion workflows for search, analytics, and generative AI, including local processing of sensitive files.
At a glance
- Language
- Python
- License
- MIT
- Stars
- 68.3k
- Forks
- 5k
- Added to OSRepos
- July 3, 2026
- Last analyzed
- October 3, 2026
Topics
Click on any tag to explore related repositories
Use at your own risk
OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of code from these repositories is the user's own responsibility. Always review the repository, source code, dependencies, licenses, and security implications before running or installing anything. OSRepos is not responsible for issues, damages, or losses resulting from third-party repositories.
Overview
Docling is a Python document parsing and conversion toolkit. It turns documents from formats such as PDF, office files, images, and audio into a unified representation that can be exported for downstream use, helping teams avoid building separate ingestion pipelines for each format.
Its PDF processing includes layout, reading order, and table understanding, while integrations connect the conversion workflow to AI and retrieval frameworks. Local execution is an option for sensitive data or environments without network access.
Key Features
- Parses a broad range of document, image, and media formats, including PDF, DOCX, XLSX, HTML, and audio.
- Analyzes PDF layout, reading order, tables, code, formulas, and images.
- Exports structured content to Markdown, HTML, WebVTT, and lossless JSON, among other formats.
- Provides OCR support for scanned documents and images.
- Offers integrations with LangChain, LlamaIndex, Haystack, and Crew AI.
- Supports a CLI and Python API, plus API-server and MCP options.
- Can run locally, including in air-gapped environments.
Use Cases
- AI application developers can convert mixed-format files into content for retrieval-augmented generation pipelines.
- Data and research teams can extract readable text and tables from PDFs and office documents for analysis.
- Organizations handling sensitive documents can run parsing locally rather than sending files to a hosted conversion service.
- Developers can build ingestion services or agent workflows using the Python API, CLI, integrations, or server options.
Project Facts
- Language: Python
- License: MIT
- Stars: 68.3k
- Forks: 5k
- Topics: ai, convert, document-parser, document-parsing, documents, docx, html, markdown, pdf, pdf-converter, pdf-to-json, pdf-to-text, pptx, tables, xlsx
- Archived: No
Getting Started
Install the package:
pip install docling
Then follow the README and documentation for conversion examples and configuration.
Alternatives
- marker: Marker converts documents to Markdown, JSON, HTML, or chunks using local models, while Docling emphasizes broad-format parsing and structured document representations.
- unstructured: Unstructured focuses on parsing and preprocessing files into elements for downstream workflows, while Docling provides document understanding and structured exports.
- opendataloader-pdf: OpenDataLoader PDF specializes in PDF parsing and tagged-PDF remediation, while Docling handles a wider range of document formats.
- xberg: Xberg is a Rust-based extraction engine with API and MCP interfaces, while Docling is a Python library centered on document understanding.
Considerations
- The README specifies Python 3.10 or higher; Python 3.9 support was dropped in version 2.70.0.
- Parsing quality and available capabilities depend on the input format and selected pipeline or models. Consult the documentation for format-specific details.
- The code is MIT licensed, but individual model licenses may differ. Review the original model packages when using them.
Source repository
Open the original repository on GitHub.
22 counted GitHub visits
Related repositories
Similar repositories that may be relevant next.

web-design: A Claude Code SKILL for Spec-First Web Page Design
October 3, 2026
The web-design project is a Claude Code SKILL designed to streamline the creation of beautiful and consistent web pages. It emphasizes a 'spec first, code second' approach, ensuring design principles are established before development begins. This tool helps generate UI, visuals, motion, and responsiveness that are consistent across pages and easily editable.

OOMWOO: Build Your Own Open-Source, Hackable Robot Vacuum Cleaner
October 2, 2026
OOMWOO is an ambitious open-source project enabling users to build their own robot vacuum cleaner using Raspberry Pi, 3D printing, and ROS2. It emphasizes local operation, hackability, and integration with Home Assistant, providing a high-quality, customizable home appliance. This project aims to deliver a fully open hardware, software, and firmware solution for autonomous home cleaning.

Shepherd: Reversible Execution Traces for Programmable Meta-Agents
October 2, 2026
Shepherd is a Python runtime substrate designed for agent work requiring inspection, reversibility, and supervision. It records agent runs as durable, inspectable execution traces, enabling meta-agents to observe, fork, replay, and revert any operation. This framework couples agents and environments using a copy-on-write fork, offering significant performance benefits and robust permission enforcement.

Agent Anvil: CI-First Evaluation Harness for Tool-Using AI Agents
October 1, 2026
Agent Anvil is a robust, CI-first evaluation harness designed for AI agents that utilize tools. It meticulously runs scenario suites, captures detailed traces of agent behavior, and provides semantic grading to identify issues. The platform excels at clustering failures and suggesting concrete fixes for prompts, tools, and guardrails, ensuring agents behave safely and effectively.