Open Source Data Extraction Tools

Discover 67 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, Library and AI. Last updated October 4, 2026.

67 repositories · updated October 4, 2026

Douyin_TikTok_Download_API: Collect Data and Download Videos

Douyin_TikTok_Download_API: Collect Data and Download Videos

A self-hosted API and console for collecting Douyin and TikTok posts, profiles, comments, and playlists, with no-watermark media downloads. It offers REST, MCP, and CLI access, and stores collected data in PostgreSQL.

PythonSelf HostedAPI
Added Nov 21, 2025 View details
feedparser: Parse RSS, Atom, and JSON Feeds in Python

feedparser: Parse RSS, Atom, and JSON Feeds in Python

feedparser is a Python library for reading RSS, Atom, and JSON feeds. It suits applications and scripts that need to retrieve and work with feed content without implementing feed parsing themselves.

PythonLibraryData Extraction
Added Nov 10, 2025 View details
marker: Convert Documents into Structured Text

marker: Convert Documents into Structured Text

Marker converts PDFs and other documents into Markdown, JSON, HTML, or chunks, preserving structure such as tables, equations, and images. It suits developers building document-processing workflows who can run its local models and inference backend.

PythonAIPDF
Added Nov 9, 2025 View details
instructor: Extract Validated Structured Data from LLMs

instructor: Extract Validated Structured Data from LLMs

Instructor turns LLM responses into validated, typed Python objects using Pydantic models. It is useful for applications that need dependable extraction across providers, with retries and streaming handled through a consistent API.

PythonLLMAI
Added Nov 8, 2025 View details
python-readability: Extract Main Text from HTML Pages

python-readability: Extract Main Text from HTML Pages

python-readability extracts an article’s main text and title from an HTML document, reducing navigation and other page clutter. It is a Python library for applications that need readable page content, with a command-line option for local files and URLs.

PythonLibraryData Extraction
Added Nov 7, 2025 View details
Argus: Gather Information for Security Reconnaissance

Argus: Gather Information for Security Reconnaissance

Argus is a Python toolkit that combines network, web application, and threat-intelligence reconnaissance modules in an interactive CLI. It suits analysts who want to run and manage varied checks from one tool, with explicit authorization for every target.

PythonCLICybersecurity
Added Nov 4, 2025 View details
mammoth.js: Convert Word Documents to Clean HTML

mammoth.js: Convert Word Documents to Clean HTML

Mammoth.js converts DOCX files into clean, semantic HTML for JavaScript applications and command-line workflows. It suits projects that need readable content rather than a pixel-perfect reproduction of Word formatting.

JavaScriptNode.jsLibrary
Added Oct 29, 2025 View details
gptpdf: Convert PDFs into Markdown with Vision Models

gptpdf: Convert PDFs into Markdown with Vision Models

gptpdf turns PDF pages into Markdown using a vision-capable model, aiming to preserve layouts such as tables, formulas, and figures. It suits developers who need an API-based PDF parsing component and can provide access to a compatible model.

PythonAILLM
Added Oct 24, 2025 View details
telegram-scraper-TeleGraphite: Collect Telegram Channel Posts as JSON

telegram-scraper-TeleGraphite: Collect Telegram Channel Posts as JSON

TeleGraphite fetches posts from public Telegram channels and saves them as JSON, with optional media downloads and contact extraction. It suits users who need a configurable, repeatable way to collect channel content for later review or processing.

PythonCLIData Extraction
Added Oct 19, 2025 View details
newspaper: Extract News Articles and Metadata in Python

newspaper: Extract News Articles and Metadata in Python

Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.

PythonLibraryWeb Scraping
Added Oct 13, 2025 View details
instructor: Extract Validated Structured Data from LLMs

instructor: Extract Validated Structured Data from LLMs

Instructor is a Python library that turns LLM responses into validated, typed data using Pydantic models. It suits applications that need reliable extraction across supported model providers without writing custom parsing and validation flows.

PythonAILLM
Added Oct 12, 2025 View details
AnyCrawl: Crawl Websites and Extract LLM-Ready Data

AnyCrawl: Crawl Websites and Extract LLM-Ready Data

AnyCrawl is a TypeScript toolkit for scraping individual pages, crawling sites, processing batches of URLs, and extracting structured data with LLMs. It suits developers building data pipelines for AI applications who need browser rendering or parallel processing.

TypeScriptNode.jsWeb Scraping
Added Oct 12, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️