Open Source Data Extraction Tools

Discover 67 open source Data Extraction repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Data Extraction projects here are most often combined with Python, Library and AI. Last updated October 4, 2026.

67 repositories · updated October 4, 2026

Ollama-OCR: Extract Text from Images and PDFs with Vision Models

Ollama-OCR: Extract Text from Images and PDFs with Vision Models

Ollama-OCR uses vision-language models served by Ollama to extract text and structured content from images and PDFs. It offers a Python package for single or batch processing and a Streamlit interface for interactive use.

PythonAIComputer Vision
Added Oct 12, 2025 View details
text-extract-api: Extract Text and Data from Documents

text-extract-api: Extract Text and Data from Documents

A self-hostable API that turns PDFs, images, and Office files into Markdown or structured JSON using OCR and Ollama models. It suits teams that need document processing and PII removal with control over where files are processed.

PythonAPIOCR
Added Oct 12, 2025 View details
sitefetch: Turn Website Content Into a Text File

sitefetch: Turn Website Content Into a Text File

sitefetch crawls a website and saves readable page content in a text file for use with AI models. It suits developers and researchers who need to collect site documentation or other web content for analysis.

TypeScriptCLIWeb Scraping
Added Oct 12, 2025 View details
maigret: Find a Person’s Accounts by Username

maigret: Find a Person’s Accounts by Username

Maigret searches thousands of websites for accounts matching a username and gathers public profile information into reports. It is an OSINT tool for investigators, researchers, and developers who need to check username reuse across platforms.

PythonCLICybersecurity
Added Oct 11, 2025 View details
Scrapling: Build Adaptive Web Scrapers and Crawlers

Scrapling: Build Adaptive Web Scrapers and Crawlers

Scrapling is a Python framework for fetching, parsing, and crawling websites, with adaptive selectors that can relocate elements after page changes. It suits projects ranging from one-off extraction to concurrent crawls and AI-assisted scraping.

PythonWeb ScrapingData Extraction
Added Oct 11, 2025 View details
pipet: Scrape and Extract Data from the Web

pipet: Scrape and Extract Data from the Web

Pipet is a command-line tool for extracting data from HTML pages, JSON endpoints, and JavaScript-rendered websites. It suits developers and technically minded users who want reusable scraping queries, shell-pipeline processing, or change monitoring.

GoCLIWeb Scraping
Added Oct 11, 2025 View details
stagehand: Build AI Agents That Use Websites

stagehand: Build AI Agents That Use Websites

Stagehand is a browser automation SDK for AI agents, combining Playwright-style controls with natural-language actions and structured data extraction. It supports TypeScript, Python, and Go, and can run with a local browser or Browserbase.

AI AgentsBrowser AutomationData Extraction
Added Oct 11, 2025 View details
Previous Page 6 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️