Open Source PDF Tools
PDF, or Portable Document Format, is a standard for sharing documents while preserving their layout across devices and operating systems. PDF tools help people create, read, edit, combine, and convert documents, as well as extract text, tables, and images. They can also process scanned pages, apply optical character recognition, and prepare document content for search or analysis.
Open source tools in this area include libraries for generating and modifying PDFs, viewers and embedding components, and utilities for parsing documents or extracting data. When choosing one, consider supported features and file types, license terms, maintenance activity, performance, security needs, and compatibility with your language or workflow. These tools are useful to developers building document features, researchers handling collections, and organizations automating document-heavy processes.
16 repositories · updated October 3, 2026

pdfminer.six: Advanced PDF Parsing and Data Extraction in Python
pdfminer.six is a powerful, community-maintained Python library designed for extracting and analyzing text data from PDF documents. It allows users to retrieve text directly from PDF source code, including details like location, font, and color. This versatile tool also supports advanced features such as CJK languages, image extraction, and various PDF specifications.

pdf-inspector: Fast Rust Library for PDF Classification and Text Extraction
pdf-inspector is a high-performance Rust library designed for intelligent PDF processing. It excels at classifying PDFs as text-based or scanned, extracting text with position awareness, and converting content to clean Markdown. This library enables smart routing decisions, significantly reducing the need for expensive OCR services for many documents.

vue-pdf-embed: A Robust PDF Embed Component for Vue 2 and Vue 3
vue-pdf-embed is a powerful and easy-to-use PDF embed component designed for Vue applications. It supports both Vue 2 and Vue 3, offering features like password-protected document handling, text and annotation layers, and no external peer dependencies. This component simplifies the integration of PDF viewing directly into your web projects.

QuestPDF: Generate PDF Documents with C#
QuestPDF is a C# library for creating structured PDF documents through a fluent, component-based API. It suits .NET developers building reports, invoices, and other data-driven documents without relying on HTML-to-PDF conversion.

opendataloader-pdf: Extract Structured Data and Accessibility Tags from PDFs
OpenDataLoader PDF parses digital, scanned, and tagged PDFs into structured formats for AI and document workflows. It also automates conversion of untagged PDFs into Tagged PDFs, with optional hybrid processing for complex documents.

Article-Assistant--RAG-Telegram-Bot: Ask Questions About Documents
A Telegram bot that turns web articles, PDFs, text files, and YouTube transcripts into searchable knowledge bases. It uses retrieval-augmented generation to answer questions with source citations and can also create summaries.

ImageToolbox: Edit, Convert, and Process Images on Android
ImageToolbox is a feature-rich Android app for photo editing, image conversion, OCR, PDF tasks, and other media workflows. It suits users who want many image utilities in one place, with distribution through Google Play, F-Droid, and GitHub releases.

Maroto: A Go Library for Creating PDFs with Bootstrap-like Layouts
Maroto is an open-source Go library designed for creating PDFs in a fast and simple manner. It draws inspiration from Bootstrap, allowing developers to structure PDF content using a familiar grid system. This tool leverages gofpdf to provide an intuitive way to generate professional documents, automatically handling page breaks and headers.

pdf-craft: Convert Scanned PDFs into Markdown and EPUB
PDF Craft uses OCR to turn scanned books and documents into editable Markdown or EPUB, with optional translation and translated PDF output. It is a Python library for developers and readers who need to process scanned material.

unstructured: Turn Documents Into Structured Data
Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

awesome-AI-books: A Curated Collection of AI and Machine Learning Resources
The awesome-AI-books repository by zslucky is a comprehensive collection of AI-related books and PDFs, designed for learning and research. It offers a wide range of resources, from introductory theory and mathematics to advanced topics like deep learning and quantum AI. This repository also includes links to various AI playground models and research organizations, making it an invaluable hub for anyone interested in artificial intelligence.

pdfplumber: Extract Text, Tables, and Objects from PDFs
pdfplumber is a Python library for inspecting machine-generated PDFs and extracting text, tables, and positioned page objects. It suits developers who need more layout detail and visual debugging than basic text extraction provides.