Open Source Web Scraping Tools
Discover 30 open source Web Scraping repositories from GitHub, each with an analysis of what it does, key features, use cases and alternatives. Web Scraping projects here are most often combined with Data Extraction, Python and TypeScript. Last updated October 3, 2026.
30 repositories · updated October 3, 2026
favicon-downloader: Find and Download Website Favicons
favicon-downloader is a TypeScript web tool for finding and downloading a site's available favicon sizes. It also provides HTML snippets and an API that retrieves icons from web pages and favicon services.

awesome-crawler: Find Web Crawling Tools Across Languages
awesome-crawler is a curated directory of web crawlers, spiders, scraping libraries, and related resources across many programming languages. Use it to compare starting points for a project rather than as a crawler implementation itself.

Scraperr: Scrape Websites Without Writing Code
Scraperr is a self-hosted web scraping application for collecting structured data from websites without writing scraping code. It supports XPath extraction, domain spidering, queued jobs, media downloads, and data export.

brightdata-mcp: Give AI Agents Access to Web Data
Bright Data MCP connects MCP-compatible agents to web search, scraping, structured extraction, and remote browser automation. It suits teams that need current public-web data without managing proxies or browser infrastructure, using a Bright Data API token.

turboseek: Build an AI Search Engine with Web Sources
TurboSeek is a TypeScript web app that answers questions using web search results and language models. It is suited to developers who want to explore or adapt a Perplexity-style search experience, with external API credentials required to run it.

PinescriptV6-docs-crawler: Crawl and Chunk Pine Script Docs
Crawl TradingView’s Pine Script v6 documentation and turn it into cleaned Markdown and heading-aware chunks for search or RAG pipelines. Incremental hashing helps avoid reprocessing pages that have not changed.

deepscrape: Scrape Websites and Extract Structured Data
DeepScrape is a self-hosted TypeScript service for scraping and crawling websites, returning clean content or structured data. It combines HTTP fetching, Playwright, optional LLM extraction, and APIs for building data pipelines and agent workflows.

ClickUi: A Desktop Assistant for Chatting with AI Models
ClickUi is a Python desktop assistant that opens with a hotkey and supports text and voice conversations with local or API-based AI models. It suits people who want AI tools available alongside their everyday computer work.

Douyin_TikTok_Download_API: Collect Data and Download Videos
A self-hosted API and console for collecting Douyin and TikTok posts, profiles, comments, and playlists, with no-watermark media downloads. It offers REST, MCP, and CLI access, and stores collected data in PostgreSQL.

python-readability: Extract Main Text from HTML Pages
python-readability extracts an article’s main text and title from an HTML document, reducing navigation and other page clutter. It is a Python library for applications that need readable page content, with a command-line option for local files and URLs.

newsnow: Read Real-Time Trending News
NewsNow is a web app for reading trending news from multiple sources in a clean interface. Use the hosted service for quick access, or deploy your own instance and connect it to an MCP client.

newspaper: Extract News Articles and Metadata in Python
Newspaper3k is a Python library for crawling news sites and extracting article text, metadata, images, keywords, and summaries. It suits developers building news aggregation, monitoring, or text-processing workflows.