Open Source Data Extraction Tools
Data extraction is the process of collecting useful information from websites, documents, feeds, and other sources, then turning it into a form that can be searched, analyzed, or used by software. It helps reduce manual copying and makes information from varied formats more consistent. Depending on the source, extraction may involve parsing page structure, handling dynamic content, recognizing text and tables, or converting records into structured formats for later use.
Open source tools in this area include web scrapers and crawlers, document and PDF parsers, feed readers, and libraries for extracting article text or metadata. When choosing one, consider supported formats and sites, output quality, reliability, maintenance activity, license, dependencies, and integration with your existing workflow. These tools are useful to developers, researchers, data teams, and organizations that need repeatable ways to gather and prepare information.
45 repositories · updated October 3, 2026

xlrd: Read Data from Legacy Excel Files
xlrd is a Python library for reading cell data and formatting from legacy Excel .xls files. It suits scripts and applications that need to inspect older spreadsheets, but it does not read newer Excel formats such as .xlsx.

pdfminer.six: Advanced PDF Parsing and Data Extraction in Python
pdfminer.six is a powerful, community-maintained Python library designed for extracting and analyzing text data from PDF documents. It allows users to retrieve text directly from PDF source code, including details like location, font, and color. This versatile tool also supports advanced features such as CJK languages, image extraction, and various PDF specifications.

awesome-skills: Community Skills for Apify Agents
A community collection of agent skills that guide AI coding assistants in using Apify Actors for tasks such as research, lead generation, and data collection. Install the collection with the skills CLI and use it with supported coding agents.

python-phonenumbers: Parse, Validate, and Format Phone Numbers
A Python port of Google’s libphonenumber for parsing, validating, and formatting international phone numbers. It also supports matching numbers in text and optional location, carrier, and timezone lookups.

Lassie: Web Content Retrieval for Humans
Lassie is a powerful Python library designed for efficient web content retrieval. It simplifies the process of extracting essential information like titles, descriptions, images, and videos from various web pages. This tool is ideal for developers needing to programmatically fetch and parse web content with ease.

toapi: Declaratively Turn Any Website into a JSON API
toapi is a powerful Python library designed to transform any website into a clean JSON API declaratively. It enables users to define desired data fields using CSS selectors, fetching and parsing web pages on demand. With built-in caching and support for dynamic content, toapi simplifies web data extraction without complex crawlers or databases.

Grab: A Powerful Python Web Scraping Framework
Grab is a robust Python web scraping framework designed to simplify complex data extraction tasks. It provides comprehensive tools for handling network requests, processing scraped content, and managing asynchronous operations through its powerful Spider component. Developers can leverage features like automatic cookie support, HTTP/SOCKS proxies, and XPath queries for efficient web data collection.

docling: Convert Documents into Structured Content
Docling converts PDFs and many other document formats into structured representations and exports such as Markdown and JSON. It is suited to developers building document ingestion workflows for search, analytics, and generative AI, including local processing of sensitive files.

llama_cloud_services: Parse Documents and Use Llama Cloud Services
A TypeScript client and service repository for Llama Cloud document processing and knowledge workflows. It is deprecated; new projects should use the replacement packages linked by the maintainers.

hiring-agent: Score Resumes with AI and GitHub Signals
Hiring Agent turns PDF resumes into structured profiles, enriches them with GitHub data, and scores them against configurable role rubrics. It is for teams exploring explainable resume prioritization, with human review remaining essential.

opendataloader-pdf: Extract Structured Data and Accessibility Tags from PDFs
OpenDataLoader PDF parses digital, scanned, and tagged PDFs into structured formats for AI and document workflows. It also automates conversion of untagged PDFs into Tagged PDFs, with optional hybrid processing for complex documents.

GLM-OCR: Recognize Text and Structure in Documents
GLM-OCR is a multimodal OCR model and SDK for extracting text and layout from complex documents. Use it through a hosted API or deploy the pipeline with supported inference servers for local control.