Open Source Data Extraction Tools

Discover 55 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, AI and Web Scraping. Last updated October 3, 2026.

55 repositories · updated October 3, 2026

Leo-Health-Core: Import Health Exports into Local SQLite

Leo-Health-Core: Import Health Exports into Local SQLite

Leo-Health-Core parses Apple Health and Whoop exports into a normalized SQLite database for local querying and dashboards. It suits people who want to explore wearable data privately with standard SQL, without sending it to a network service.

PythonDatabasePrivacy
Added May 17, 2026 View details
waymore: Find and Download Archived URLs

waymore: Find and Download Archived URLs

waymore collects historical URLs from web archives and threat-intelligence sources, and can download archived responses for further analysis. It is aimed at security researchers and bug bounty hunters who value broad coverage over speed.

PythonCLISecurity
Added May 15, 2026 View details
firecrawl: Search, Scrape, and Crawl the Web for AI

firecrawl: Search, Scrape, and Crawl the Web for AI

Firecrawl turns web pages and sites into content that AI applications can use, including Markdown and structured data. It suits developers building agents or data pipelines that need web search, extraction, and crawling through an API or self-hosted deployment.

TypeScriptAIAI Agents
Added May 13, 2026 View details
trafilatura: Extract Text and Metadata from Web Pages

trafilatura: Extract Text and Metadata from Web Pages

Trafilatura is a Python library and CLI for discovering web content and extracting readable text and metadata from HTML. It suits developers and researchers building crawlers, web corpora, or content-processing pipelines that need structured output.

PythonWeb ScrapingData Extraction
Added May 1, 2026 View details
index: Automate Complex Tasks in a Web Browser

index: Automate Complex Tasks in a Web Browser

Index is a Python browser agent that uses reasoning-capable language models to carry out multi-step web tasks and return structured results. It offers a CLI, a Python interface, and a serverless API, but the repository is archived.

PythonAILLM
Added Apr 27, 2026 View details
OmniParse: Turn Documents and Media into AI-Ready Data

OmniParse: Turn Documents and Media into AI-Ready Data

OmniParse is a self-hosted Python service that converts documents, images, audio, video, and web pages into structured output for GenAI workflows. It suits teams building ingestion pipelines that want local parsing, but requires Linux and a GPU with at least 8–10 GB of VRAM.

PythonAIAPI
Added Apr 7, 2026 View details
notebooks: Learn and Apply Computer Vision Models

notebooks: Learn and Apply Computer Vision Models

Roboflow notebooks is a hands-on tutorial collection for computer vision, covering model training, inference, detection, segmentation, and related tasks. Use it to explore techniques and run examples in hosted notebook environments.

Computer VisionMachine LearningDeep Learning
Added Apr 6, 2026 View details
Docling: Streamlining Document Processing for Generative AI

Docling: Streamlining Document Processing for Generative AI

Docling is a powerful Python library designed to simplify document processing and prepare diverse formats for generative AI applications. It excels at parsing various document types, including advanced PDF understanding, and offers seamless integrations with popular AI frameworks. With Docling, developers can efficiently extract, transform, and utilize document content for their AI models.

PythonAIDocument Parsing
Added Mar 22, 2026 View details
NeoStumbler: Map Wireless Networks for Geolocation

NeoStumbler: Map Wireless Networks for Geolocation

NeoStumbler is an Android app for recording cell tower, Wi-Fi and Bluetooth beacon locations and contributing them to compatible geolocation services. It suits people who want to help improve coverage data or inspect and export their own observations.

AndroidGeolocationNetworking
Added Mar 21, 2026 View details
PaddleOCR: Extract Text and Structure from Images and PDFs

PaddleOCR: Extract Text and Structure from Images and PDFs

PaddleOCR is a Python toolkit for recognizing text and parsing document layouts in images and PDFs. It suits developers building OCR, document-processing, and retrieval workflows who need structured output and multilingual recognition.

PythonOCRComputer Vision
Added Mar 14, 2026 View details
DiscordChatExporter: Save Discord Chats to Files

DiscordChatExporter: Save Discord Chats to Files

DiscordChatExporter exports message history from Discord direct messages, group chats, and server channels. It suits people who need portable, offline chat archives in formats such as HTML, TXT, CSV, or JSON.

CsharpDesktop AppCLI
Added Mar 13, 2026 View details
pdf-craft: Convert Scanned PDFs into Markdown and EPUB

pdf-craft: Convert Scanned PDFs into Markdown and EPUB

PDF Craft uses OCR to turn scanned books and documents into editable Markdown or EPUB, with optional translation and translated PDF output. It is a Python library for developers and readers who need to process scanned material.

PythonOCRPDF
Added Mar 9, 2026 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️