Open Source Data Extraction Tools
Discover 66 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, Library and AI. Last updated October 4, 2026.
66 repositories · updated October 4, 2026

wallabag: Save Web Pages for Later Reading
wallabag is a self-hostable web application for saving, classifying, and reading articles later. It extracts page content to provide a less cluttered reading experience, with companion apps and a browser extension available.

brightdata-mcp: Give AI Agents Access to Web Data
Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

YTSage: Download YouTube Videos with a Desktop App
YTSage is a cross-platform Python desktop downloader built around yt-dlp. Its PySide6 interface helps users download video, audio, subtitles, and playlists, with options such as SponsorBlock and format selection.

xberg: Extract Text and Structure from Documents
Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs
Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

markitdown: Convert Documents and Files to Markdown
MarkItDown converts documents and other files into Markdown for LLM and text-analysis workflows. It offers a Python library and command-line interface, with optional format-specific dependencies and plugins.

pypdf: Read and Manipulate PDF Files in Python
pypdf is a pure-Python library for reading PDFs and changing their pages or document data. It suits Python developers who need PDF operations inside scripts or applications without relying on a separate command-line tool.

e2m: Convert Documents and Media into Markdown
E2M is a Python library for parsing documents, web pages, and audio into Markdown through configurable parser and converter components. It is aimed at developers preparing varied source material for RAG, training data, or downstream text workflows.

deepscrape: Scrape Websites and Extract Structured Data
DeepScrape is a self-hosted TypeScript service for scraping and crawling websites, returning clean content or structured data. It combines HTTP fetching, Playwright, optional LLM extraction, and APIs for building data pipelines and agent workflows.

graphrag: Build Knowledge Graphs for LLM Question Answering
Microsoft GraphRAG is a Python pipeline that uses LLMs to turn unstructured text into structured, graph-based context for question answering. It is suited to teams exploring graph-enhanced retrieval over private data, with indexing costs and maintenance-mode status to consider.

dlt: Load Data from Sources into Analytics Destinations
dlt is a Python library for building data-loading pipelines that extract data from APIs, databases, files, and Python objects, then load it into analytics destinations. It handles schema inference, normalization, and incremental loading within your existing code.

attachments: Turn Files Into LLM-Ready Context
attachments is a Python library and CLI that turns documents, images, audio, and other inputs into structured text and image artifacts for LLM workflows. It supports local processing, optional service fallback, and adapters for prompts, chat APIs, and RAG chunks.