Open Source Web Scraping Tools

Discover 30 open source Web Scraping repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Web Scraping projects here are most often combined with Data Extraction, Python and TypeScript. Last updated October 3, 2026.

30 repositories · updated October 3, 2026

favicon-downloader: Find and Download Website Favicons

favicon-downloader: Find and Download Website Favicons

favicon-downloader is a TypeScript web tool for finding and downloading a site's available favicon sizes. It also provides HTML snippets and an API that retrieves icons from web pages and favicon services.

TypeScriptWeb AppDeveloper Tools
Added Mar 21, 2026 View details
awesome-crawler: Find Web Crawling Tools Across Languages

awesome-crawler: Find Web Crawling Tools Across Languages

awesome-crawler is a curated directory of web crawlers, spiders, scraping libraries, and related resources across many programming languages. Use it to compare starting points for a project rather than as a crawler implementation itself.

Awesome ListWeb ScrapingResources
Added Mar 1, 2026 View details
Scraperr: Scrape Websites Without Writing Code

Scraperr: Scrape Websites Without Writing Code

Scraperr is a self-hosted web scraping application for collecting structured data from websites without writing scraping code. It supports XPath extraction, domain spidering, queued jobs, media downloads, and data export.

Web ScrapingData ExtractionSelf Hosted
Added Feb 16, 2026 View details
brightdata-mcp: Give AI Agents Access to Web Data

brightdata-mcp: Give AI Agents Access to Web Data

Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

JavaScriptAI AgentsMCP
Added Jan 22, 2026 View details
turboseek: Build an AI Search Engine with Web Sources

turboseek: Build an AI Search Engine with Web Sources

TurboSeek is a TypeScript web app that answers questions using web search results and language models. It is suited to developers who want to explore or adapt a Perplexity-style search experience, with external API credentials required to run it.

TypeScriptAILLM
Added Jan 17, 2026 View details
PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs

Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

PythonWeb ScrapingDocumentation
Added Dec 27, 2025 View details
deepscrape: Scrape Websites and Extract Structured Data

deepscrape: Scrape Websites and Extract Structured Data

DeepScrape is a self-hosted TypeScript service for scraping and crawling websites, returning clean content or structured data. It combines HTTP fetching, Playwright, optional LLM extraction, and APIs for building data pipelines and agent workflows.

TypeScriptWeb ScrapingData Extraction
Added Dec 19, 2025 View details
ClickUi: A Desktop Assistant for Chatting with AI Models

ClickUi: A Desktop Assistant for Chatting with AI Models

ClickUi is a Python desktop assistant that opens with a hotkey and supports text and voice conversations with local or API-based AI models. It suits people who want AI tools available alongside their everyday computer work.

PythonAIDesktop App
Added Dec 9, 2025 View details
Douyin_TikTok_Download_API: Collect Data and Download Videos

Douyin_TikTok_Download_API: Collect Data and Download Videos

A self-hosted API and console for collecting Douyin and TikTok posts, profiles, comments, and playlists, with no-watermark media downloads. It offers REST, MCP, and CLI access, and stores collected data in PostgreSQL.

PythonSelf HostedAPI
Added Nov 21, 2025 View details
python-readability: Extract Main Text from HTML Pages

python-readability: Extract Main Text from HTML Pages

python-readability extracts an article’s main text and title from an HTML document, reducing navigation and other page clutter. It is a Python library for applications that need readable page content, with a command-line option for local files and URLs.

PythonLibraryData Extraction
Added Nov 7, 2025 View details
newsnow: Read Real-Time Trending News

newsnow: Read Real-Time Trending News

NewsNow is a web app for reading trending news from multiple sources in a clean interface. Use the hosted service for quick access, or deploy your own instance and connect it to an MCP client.

TypeScriptWeb AppReal Time
Added Nov 6, 2025 View details
newspaper: Extract News Articles and Metadata in Python

newspaper: Extract News Articles and Metadata in Python

Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.

PythonLibraryWeb Scraping
Added Oct 13, 2025 View details

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️