Open Source Data Extraction Tools
Discover 55 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, AI and Web Scraping. Last updated October 3, 2026.
55 repositories · updated October 3, 2026

Leo-Health-Core: Import Health Exports into Local SQLite
Leo-Health-Core parses Apple Health and Whoop exports into a normalized SQLite database for local querying and dashboards. It suits people who want to explore wearable data privately with standard SQL, without sending it to a network service.

waymore: Find and Download Archived URLs
waymore collects historical URLs from web archives and threat-intelligence sources, and can download archived responses for further analysis. It is aimed at security researchers and bug bounty hunters who value broad coverage over speed.

firecrawl: Search, Scrape, and Crawl the Web for AI
Firecrawl turns web pages and sites into content that AI applications can use, including Markdown and structured data. It suits developers building agents or data pipelines that need web search, extraction, and crawling through an API or self-hosted deployment.

trafilatura: Extract Text and Metadata from Web Pages
Trafilatura is a Python library and CLI for discovering web content and extracting readable text and metadata from HTML. It suits developers and researchers building crawlers, web corpora, or content-processing pipelines that need structured output.

index: Automate Complex Tasks in a Web Browser
Index is a Python browser agent that uses reasoning-capable language models to carry out multi-step web tasks and return structured results. It offers a CLI, a Python interface, and a serverless API, but the repository is archived.

OmniParse: Turn Documents and Media into AI-Ready Data
OmniParse is a self-hosted Python service that converts documents, images, audio, video, and web pages into structured output for GenAI workflows. It suits teams building ingestion pipelines that want local parsing, but requires Linux and a GPU with at least 8–10 GB of VRAM.

notebooks: Learn and Apply Computer Vision Models
Roboflow notebooks is a hands-on tutorial collection for computer vision, covering model training, inference, detection, segmentation, and related tasks. Use it to explore techniques and run examples in hosted notebook environments.

Docling: Streamlining Document Processing for Generative AI
Docling is a powerful Python library designed to simplify document processing and prepare diverse formats for generative AI applications. It excels at parsing various document types, including advanced PDF understanding, and offers seamless integrations with popular AI frameworks. With Docling, developers can efficiently extract, transform, and utilize document content for their AI models.

NeoStumbler: Map Wireless Networks for Geolocation
NeoStumbler is an Android app for recording cell tower, Wi-Fi and Bluetooth beacon locations and contributing them to compatible geolocation services. It suits people who want to help improve coverage data or inspect and export their own observations.

PaddleOCR: Extract Text and Structure from Images and PDFs
PaddleOCR is a Python toolkit for recognizing text and parsing document layouts in images and PDFs. It suits developers building OCR, document-processing, and retrieval workflows who need structured output and multilingual recognition.

DiscordChatExporter: Save Discord Chats to Files
DiscordChatExporter exports message history from Discord direct messages, group chats, and server channels. It suits people who need portable, offline chat archives in formats such as HTML, TXT, CSV, or JSON.

pdf-craft: Convert Scanned PDFs into Markdown and EPUB
PDF Craft uses OCR to turn scanned books and documents into editable Markdown or EPUB, with optional translation and translated PDF output. It is a Python library for developers and readers who need to process scanned material.