Open Source Web Scraping Tools
Web scraping is the automated collection of information from websites. It can turn pages, tables, and dynamically loaded content into structured data for research, monitoring, analysis, and other applications. Scraping tools handle tasks such as fetching pages, rendering JavaScript, selecting relevant content, and managing data output, reducing the effort of gathering information by hand. Responsible use includes respecting site terms, privacy requirements, and limits on automated access.
Open source options range from lightweight libraries and command-line utilities to browser automation frameworks, full crawlers, and services that prepare web content for AI systems. When choosing a tool, consider its license, maintenance activity, documentation, language and runtime requirements, and ability to handle the sites and data formats you need. Also assess reliability, scalability, and integration with your existing workflows. These tools are useful to developers, researchers, analysts, and organizations that need repeatable ways to collect public web data.
29 repositories · updated October 3, 2026

browser-rs-mcp: Lightweight Stealth Browser Server for AI Agents in Rust
browser-rs-mcp is a lightweight, stealth-oriented Model Context Protocol (MCP) browser server written in Rust. It enables multiple AI agents to share a single, persistent Chrome instance, offering over 64 Playwright-style tools without requiring a Node.js runtime. This project is ideal for parallel web scraping, automation, and QA, providing efficient and isolated tab control for each agent.

Lassie: Web Content Retrieval for Humans
Lassie is a powerful Python library designed for efficient web content retrieval. It simplifies the process of extracting essential information like titles, descriptions, images, and videos from various web pages. This tool is ideal for developers needing to programmatically fetch and parse web content with ease.

requests-html: Pythonic HTML Parsing with JavaScript Support
requests-html is a Python library designed to simplify HTML parsing and web scraping. It extends the familiar Requests experience with powerful parsing capabilities, including full JavaScript support via Chromium, CSS selectors, and XPath. This makes it an ideal tool for developers needing to interact with dynamic web content.

toapi: Declaratively Turn Any Website into a JSON API
toapi is a powerful Python library designed to transform any website into a clean JSON API declaratively. It enables users to define desired data fields using CSS selectors, fetching and parsing web pages on demand. With built-in caching and support for dynamic content, toapi simplifies web data extraction without complex crawlers or databases.

Grab: A Powerful Python Web Scraping Framework
Grab is a robust Python web scraping framework designed to simplify complex data extraction tasks. It provides comprehensive tools for handling network requests, processing scraped content, and managing asynchronous operations through its powerful Spider component. Developers can leverage features like automatic cookie support, HTTP/SOCKS proxies, and XPath queries for efficient web data collection.

MechanicalSoup: A Python Library for Automating Website Interaction
MechanicalSoup is a powerful Python library designed for automating interactions with websites. Built upon Requests and BeautifulSoup, it simplifies tasks like storing cookies, following redirects, and submitting forms. It's an excellent tool for web automation tasks that don't require JavaScript execution.

proxy-list: A Daily Updated List of Free Proxy Servers
The `clarketm/proxy-list` repository provides a comprehensive, daily updated collection of free, public, forward proxy servers. It offers various formats for easy access, including raw IP:PORT lists and detailed information on country, anonymity, and type. This resource is invaluable for developers and users needing reliable proxy access for various networking tasks.

Agent-Reach: Give AI Agents Access to Web Sources
Agent-Reach is a Python CLI that helps AI agents read and search across websites and platforms such as YouTube, Reddit, GitHub, and Twitter. It installs and checks integrations, with some sources available immediately and others requiring login or configuration.

CloakBrowser: Automate Chromium with Modified Fingerprints
CloakBrowser wraps a custom Chromium build with source-level fingerprint changes and Playwright- and Puppeteer-compatible APIs. It is aimed at browser automation and web scraping where standard browser automation is frequently flagged by bot detection.

ai-website-cloner-template: Recreate Websites as Next.js Apps
A template that equips AI coding agents to inspect a website and recreate it as a Next.js application. It suits developers rebuilding sites they own, recovering lost source code, or studying real-world interface implementations.

firecrawl: Search, Scrape, and Crawl the Web for AI
Firecrawl turns web pages and sites into content that AI applications can use, including Markdown and structured data. It suits developers building agents or data pipelines that need web search, extraction, and crawling through an API or self-hosted deployment.

Cheerio: Fast and Flexible HTML/XML Parsing and Manipulation Library
Cheerio is a popular library for parsing and manipulating HTML and XML documents in Node.js. It provides a jQuery-like API, making it easy to select, traverse, and modify elements with proven syntax. Known for its blazingly fast performance and incredible flexibility, Cheerio is an excellent choice for web scraping and server-side DOM manipulation.