Open Source PDF Tools
Discover 21 open source PDF repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. PDF projects here are most often combined with Python, OCR and Data Extraction. Last updated October 3, 2026.
21 repositories · updated October 3, 2026

pdf-craft: Convert Scanned PDFs into Markdown and EPUB
PDF Craft uses OCR to turn scanned books and documents into editable Markdown or EPUB, with optional translation and translated PDF output. It is a Python library for developers and readers who need to process scanned material.

unstructured: Turn Documents Into Structured Data
Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

pdfplumber: Extract Text, Tables, and Objects from PDFs
pdfplumber is a Python library for inspecting machine-generated PDFs and extracting text, tables, and positioned page objects. It suits developers who need more layout detail and visual debugging than basic text extraction provides.

papermerge: Organize and Search Scanned Documents
Papermerge is a web-based document management system for organizing scanned archives. It uses OCR and full-text search to make documents easier to find, but this repository is archived and development has moved to papermerge-core.

pypdf: A Powerful Pure-Python Library for PDF Manipulation
pypdf is a free and open-source pure-Python library designed for comprehensive PDF manipulation. It allows users to split, merge, crop, and transform PDF pages, as well as add custom data, viewing options, and passwords. The library also supports extracting text and metadata from PDF files, making it a versatile tool for various PDF-related tasks.

open-notebooklm: Turn PDFs Into Podcast Audio
Open NotebookLM turns a PDF into an AI-generated podcast dialogue and MP3. It suits readers who want an audio-style overview of a document, and requires a Fireworks API key to run.

paperlib: Manage and Organize Academic Papers
Paperlib is a cross-platform desktop tool for collecting, searching, and organizing academic papers. It helps researchers find and correct publication metadata, manage PDFs and notes, and export references while writing.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

text-extract-api: Advanced Document Extraction, OCR, and PII Removal with LLMs
text-extract-api is a powerful API designed for extracting and parsing text from various document formats, including PDF, Word, and PPTX. It utilizes modern OCRs and Ollama-supported LLMs for highly accurate text extraction, PII removal, and conversion to structured JSON or Markdown, all while maintaining data privacy through its self-hosted architecture.