Open Source Data Extraction Tools

Discover 57 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, AI and Web Scraping. Last updated October 3, 2026.

57 repositories · updated October 3, 2026

maestro: Fine-Tune Vision-Language Models

maestro: Fine-Tune Vision-Language Models

Maestro streamlines fine-tuning for multimodal vision-language models with ready-to-use recipes, a CLI, and a Python API. It suits developers adapting supported models to tasks such as JSON extraction and object detection.

PythonMachine LearningComputer Vision
Added Mar 2, 2026 View details
awesome-crawler: Find Web Crawling Tools Across Languages

awesome-crawler: Find Web Crawling Tools Across Languages

awesome-crawler is a curated directory of web crawlers, spiders, scraping libraries, and related resources across many programming languages. Use it to compare starting points for a project rather than as a crawler implementation itself.

Awesome ListWeb ScrapingResources
Added Mar 1, 2026 View details
Scraperr: Scrape Websites Without Writing Code

Scraperr: Scrape Websites Without Writing Code

Scraperr is a self-hosted web scraping application for collecting structured data from websites without writing scraping code. It supports XPath extraction, domain spidering, queued jobs, media downloads, and data export.

Web ScrapingData ExtractionSelf Hosted
Added Feb 16, 2026 View details
unstructured: Turn Documents Into Structured Data

unstructured: Turn Documents Into Structured Data

Unstructured is a Python library for parsing and preprocessing documents into structured elements for downstream applications, including LLM workflows. It supports many file types, with format-specific dependencies for some inputs.

PythonLLMData Extraction
Added Feb 10, 2026 View details
RAG-Anything: Search Documents Across Text, Images, Tables, and Equations

RAG-Anything: Search Documents Across Text, Images, Tables, and Equations

RAG-Anything extends LightRAG with a pipeline for parsing and querying multimodal documents. It is aimed at teams and researchers who need one retrieval system for mixed-content files rather than text-only RAG.

PythonRAGFramework
Added Jan 31, 2026 View details
pdfplumber: Extract Text, Tables, and Objects from PDFs

pdfplumber: Extract Text, Tables, and Objects from PDFs

pdfplumber is a Python library for inspecting machine-generated PDFs and extracting text, tables, and positioned page objects. It suits developers who need more layout detail and visual debugging than basic text extraction provides.

PythonLibraryData Extraction
Added Jan 24, 2026 View details
wallabag: Save Web Pages for Later Reading

wallabag: Save Web Pages for Later Reading

wallabag is a self-hostable web application for saving, classifying, and reading articles later. It extracts page content to provide a less cluttered reading experience, with companion apps and a browser extension available.

PHPSelf HostedProductivity
Added Jan 23, 2026 View details
brightdata-mcp: Give AI Agents Access to Web Data

brightdata-mcp: Give AI Agents Access to Web Data

Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

JavaScriptAI AgentsMCP
Added Jan 22, 2026 View details
YTSage: Download YouTube Videos with a Desktop App

YTSage: Download YouTube Videos with a Desktop App

YTSage is a cross-platform Python desktop downloader built around yt-dlp. Its PySide6 interface helps users download video, audio, subtitles, and playlists, with options such as SponsorBlock and format selection.

PythonDesktop AppCross Platform
Added Dec 31, 2025 View details
xberg: Extract Text and Structure from Documents

xberg: Extract Text and Structure from Documents

Xberg is a Rust-based document intelligence engine that extracts text, tables, metadata, and structured data from many file types. Use it as a library, CLI, REST API, or MCP server, with bindings for multiple languages.

RustData ExtractionLibrary
Added Dec 30, 2025 View details
PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

PythonWeb ScrapingDocumentation
Added Dec 27, 2025 View details
pypdf: Read and Manipulate PDF Files in Python

pypdf: Read and Manipulate PDF Files in Python

pypdf is a pure-Python library for reading PDFs and changing their pages or document data. It suits Python developers who need PDF operations inside scripts or applications without relying on a separate command-line tool.

PythonLibraryPDF
Added Dec 24, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️