Open Source Data Extraction Tools
Discover 57 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, AI and Web Scraping. Last updated October 3, 2026.
57 repositories · updated October 3, 2026

maestro: Fine-Tune Vision-Language Models
Maestro streamlines fine-tuning for multimodal vision-language models with ready-to-use recipes, a CLI, and a Python API. It suits developers adapting supported models to tasks such as JSON extraction and object detection.

awesome-crawler: Find Web Crawling Tools Across Languages
awesome-crawler is a curated directory of web crawlers, spiders, scraping libraries, and related resources across many programming languages. Use it to compare starting points for a project rather than as a crawler implementation itself.

Scraperr: Scrape Websites Without Writing Code
Scraperr is a self-hosted web scraping application for collecting structured data from websites without writing scraping code. It supports XPath extraction, domain spidering, queued jobs, media downloads, and data export.

unstructured: Turn Documents Into Structured Data
Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

RAG-Anything: Search Documents Across Text, Images, Tables, and Equations
RAG-Anything extends LightRAG with a pipeline for parsing and querying multimodal documents. It is aimed at teams and researchers who need one retrieval system for mixed-content files rather than text-only RAG.

pdfplumber: Extract Text, Tables, and Objects from PDFs
pdfplumber is a Python library for inspecting machine-generated PDFs and extracting text, tables, and positioned page objects. It suits developers who need more layout detail and visual debugging than basic text extraction provides.

wallabag: Save Web Pages for Later Reading
wallabag is a self-hostable web application for saving, classifying, and reading articles later. It extracts page content to provide a less cluttered reading experience, with companion apps and a browser extension available.

brightdata-mcp: Give AI Agents Access to Web Data
Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

YTSage: Download YouTube Videos with a Desktop App
YTSage is a cross-platform Python desktop downloader built around yt-dlp. Its PySide6 interface helps users download video, audio, subtitles, and playlists, with options such as SponsorBlock and format selection.

xberg: Extract Text and Structure from Documents
Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs
Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

pypdf: Read and Manipulate PDF Files in Python
pypdf is a pure-Python library for reading PDFs and changing their pages or document data. It suits Python developers who need PDF operations inside scripts or applications without relying on a separate command-line tool.