Open Source Data Extraction Tools
Discover 67 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, Library and AI. Last updated October 4, 2026.
67 repositories · updated October 4, 2026

Douyin_TikTok_Download_API: Collect Data and Download Videos
A self-hosted API and console for collecting Douyin and TikTok posts, profiles, comments, and playlists, with no-watermark media downloads. It offers REST, MCP, and CLI access, and stores collected data in PostgreSQL.

feedparser: Parse RSS, Atom, and JSON Feeds in Python
feedparser is a Python library for reading RSS, Atom, and JSON feeds. It suits applications and scripts that need to retrieve and work with feed content without implementing feed parsing themselves.

marker: Convert Documents into Structured Text
Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

instructor: Extract Validated Structured Data from LLMs
Instructor turns LLM responses into validated, typed Python objects using Pydantic models. It is useful for applications that need dependable extraction across providers, with retries and streaming handled through a consistent API.

python-readability: Extract Main Text from HTML Pages
python-readability extracts an article’s main text and title from an HTML document, reducing navigation and other page clutter. It is a Python library for applications that need readable page content, with a command-line option for local files and URLs.

Argus: Gather Information for Security Reconnaissance
Argus is a Python toolkit that combines network, web application, and threat-intelligence reconnaissance modules in an interactive CLI. It suits analysts who want to run and manage varied checks from one tool, with explicit authorization for every target.

mammoth.js: Convert Word Documents to Clean HTML
Mammoth.js converts DOCX files into clean, semantic HTML for JavaScript applications and command-line workflows. It suits projects that need readable content rather than a pixel-perfect reproduction of Word formatting.

gptpdf: Convert PDFs into Markdown with Vision Models
gptpdf turns PDF pages into Markdown using a vision-capable model, aiming to preserve layouts such as tables, formulas, and figures. It suits developers who need an API-based PDF parsing component and can provide access to a compatible model.

telegram-scraper-TeleGraphite: Collect Telegram Channel Posts as JSON
TeleGraphite fetches posts from public Telegram channels and saves them as JSON, with optional media downloads and contact extraction. It suits users who need a configurable, repeatable way to collect channel content for later review or processing.

newspaper: Extract News Articles and Metadata in Python
Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.

instructor: Extract Validated Structured Data from LLMs
Instructor is a Python library that turns LLM responses into validated, typed data using Pydantic models. It suits applications that need reliable extraction across supported model providers without writing custom parsing and validation flows.

AnyCrawl: Crawl Websites and Extract LLM-Ready Data
AnyCrawl is a TypeScript toolkit for scraping individual pages, crawling sites, processing batches of URLs, and extracting structured data with LLMs. It suits developers building data pipelines for AI applications who need browser rendering or parallel processing.