Open Source Data Extraction Tools

Discover 66 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, Library and AI. Last updated October 4, 2026.

66 repositories · updated October 4, 2026

wallabag: Save Web Pages for Later Reading

wallabag: Save Web Pages for Later Reading

wallabag is a self-hostable web application for saving, classifying, and reading articles later. It extracts page content to provide a less cluttered reading experience, with companion apps and a browser extension available.

PHPSelf HostedProductivity
Added Jan 23, 2026 View details
brightdata-mcp: Give AI Agents Access to Web Data

brightdata-mcp: Give AI Agents Access to Web Data

Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

JavaScriptAI AgentsMCP
Added Jan 22, 2026 View details
YTSage: Download YouTube Videos with a Desktop App

YTSage: Download YouTube Videos with a Desktop App

YTSage is a cross-platform Python desktop downloader built around yt-dlp. Its PySide6 interface helps users download video, audio, subtitles, and playlists, with options such as SponsorBlock and format selection.

PythonDesktop AppCross Platform
Added Dec 31, 2025 View details
xberg: Extract Text and Structure from Documents

xberg: Extract Text and Structure from Documents

Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

RustData ExtractionLibrary
Added Dec 30, 2025 View details
PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

PythonWeb ScrapingDocumentation
Added Dec 27, 2025 View details
markitdown: Convert Documents and Files to Markdown

markitdown: Convert Documents and Files to Markdown

MarkItDown converts documents and other files into Markdown for LLM and text-analysis workflows. It offers a Python library and command-line interface, with optional format-specific dependencies and plugins.

PythonLibraryCLI
Added Dec 27, 2025 View details
pypdf: Read and Manipulate PDF Files in Python

pypdf: Read and Manipulate PDF Files in Python

pypdf is a pure-Python library for reading PDFs and changing their pages or document data. It suits Python developers who need PDF operations inside scripts or applications without relying on a separate command-line tool.

PythonLibraryPDF
Added Dec 24, 2025 View details
e2m: Convert Documents and Media into Markdown

e2m: Convert Documents and Media into Markdown

E2M is a Python library for parsing documents, web pages, and audio into Markdown through configurable parser and converter components. It is aimed at developers preparing varied source material for RAG, training data, or downstream text workflows.

PythonLibraryAI
Added Dec 24, 2025 View details
deepscrape: Scrape Websites and Extract Structured Data

deepscrape: Scrape Websites and Extract Structured Data

DeepScrape is a self-hosted TypeScript service for scraping and crawling websites, returning clean content or structured data. It combines HTTP fetching, Playwright, optional LLM extraction, and APIs for building data pipelines and agent workflows.

TypeScriptWeb ScrapingData Extraction
Added Dec 19, 2025 View details
graphrag: Build Knowledge Graphs for LLM Question Answering

graphrag: Build Knowledge Graphs for LLM Question Answering

Microsoft GraphRAG is a Python pipeline that uses LLMs to turn unstructured text into structured, graph-based context for question answering. It is suited to teams exploring graph-enhanced retrieval over private data, with indexing costs and maintenance-mode status to consider.

PythonLLMRAG
Added Dec 17, 2025 View details
dlt: Load Data from Sources into Analytics Destinations

dlt: Load Data from Sources into Analytics Destinations

dlt is a Python library for building data-loading pipelines that extract data from APIs, databases, files, and Python objects, then load it into analytics destinations. It handles schema inference, normalization, and incremental loading within your existing code.

PythonLibraryData Extraction
Added Dec 15, 2025 View details
attachments: Turn Files Into LLM-Ready Context

attachments: Turn Files Into LLM-Ready Context

attachments is a Python library and CLI that turns documents, images, audio, and other inputs into structured text and image artifacts for LLM workflows. It supports local processing, optional service fallback, and adapters for prompts, chat APIs, and RAG chunks.

PythonLLMAI
Added Nov 24, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️