Open Source Web Scraping Tools

Web scraping is the automated collection of information from websites. It can turn pages, tables, and dynamically loaded content into structured data for research, monitoring, analysis, and other applications. Scraping tools handle tasks such as fetching pages, rendering JavaScript, selecting relevant content, and managing data output, reducing the effort of gathering information by hand. Responsible use includes respecting site terms, privacy requirements, and limits on automated access.

Open source options range from lightweight libraries and command-line utilities to browser automation frameworks, full crawlers, and services that prepare web content for AI systems. When choosing a tool, consider its license, maintenance activity, documentation, language and runtime requirements, and ability to handle the sites and data formats you need. Also assess reliability, scalability, and integration with your existing workflows. These tools are useful to developers, researchers, analysts, and organizations that need repeatable ways to collect public web data.

29 repositories · updated October 3, 2026

browser-rs-mcp: Lightweight Stealth Browser Server for AI Agents in Rust

browser-rs-mcp: Lightweight Stealth Browser Server for AI Agents in Rust

browser-rs-mcp is a lightweight, stealth-oriented Model Context Protocol (MCP) browser server written in Rust. It enables multiple AI agents to share a single, persistent Chrome instance, offering over 64 Playwright-style tools without requiring a Node.js runtime. This project is ideal for parallel web scraping, automation, and QA, providing efficient and isolated tab control for each agent.

RustAI AgentsBrowser Automation
Added Sep 25, 2026 View details
Lassie: Web Content Retrieval for Humans

Lassie: Web Content Retrieval for Humans

Lassie is a powerful Python library designed for efficient web content retrieval. It simplifies the process of extracting essential information like titles, descriptions, images, and videos from various web pages. This tool is ideal for developers needing to programmatically fetch and parse web content with ease.

ContentMetaOembed
Added Jul 25, 2026 View details
requests-html: Pythonic HTML Parsing with JavaScript Support

requests-html: Pythonic HTML Parsing with JavaScript Support

requests-html is a Python library designed to simplify HTML parsing and web scraping. It extends the familiar Requests experience with powerful parsing capabilities, including full JavaScript support via Chromium, CSS selectors, and XPath. This makes it an ideal tool for developers needing to interact with dynamic web content.

PythonWeb ScrapingHTML Parsing
Added Jul 25, 2026 View details
toapi: Declaratively Turn Any Website into a JSON API

toapi: Declaratively Turn Any Website into a JSON API

toapi is a powerful Python library designed to transform any website into a clean JSON API declaratively. It enables users to define desired data fields using CSS selectors, fetching and parsing web pages on demand. With built-in caching and support for dynamic content, toapi simplifies web data extraction without complex crawlers or databases.

APIPythonWeb Scraping
Added Jul 24, 2026 View details
Grab: A Powerful Python Web Scraping Framework

Grab: A Powerful Python Web Scraping Framework

Grab is a robust Python web scraping framework designed to simplify complex data extraction tasks. It provides comprehensive tools for handling network requests, processing scraped content, and managing asynchronous operations through its powerful Spider component. Developers can leverage features like automatic cookie support, HTTP/SOCKS proxies, and XPath queries for efficient web data collection.

PythonWeb ScrapingFramework
Added Jul 24, 2026 View details
MechanicalSoup: A Python Library for Automating Website Interaction

MechanicalSoup: A Python Library for Automating Website Interaction

MechanicalSoup is a powerful Python library designed for automating interactions with websites. Built upon Requests and BeautifulSoup, it simplifies tasks like storing cookies, following redirects, and submitting forms. It's an excellent tool for web automation tasks that don't require JavaScript execution.

MechanicalsoupPythonWeb Automation
Added Jul 24, 2026 View details
proxy-list: A Daily Updated List of Free Proxy Servers

proxy-list: A Daily Updated List of Free Proxy Servers

The `clarketm/proxy-list` repository provides a comprehensive, daily updated collection of free, public, forward proxy servers. It offers various formats for easy access, including raw IP:PORT lists and detailed information on country, anonymity, and type. This resource is invaluable for developers and users needing reliable proxy access for various networking tasks.

ProxyProxy ListProxy Server
Added Jul 8, 2026 View details
Agent-Reach: Give AI Agents Access to Web Sources

Agent-Reach: Give AI Agents Access to Web Sources

Agent-Reach is a Python CLI that helps AI agents read and search across websites and platforms such as YouTube, Reddit, GitHub, and Twitter. It installs and checks integrations, with some sources available immediately and others requiring login or configuration.

PythonAI AgentsCLI
Added Jun 21, 2026 View details
CloakBrowser: Automate Chromium with Modified Fingerprints

CloakBrowser: Automate Chromium with Modified Fingerprints

CloakBrowser wraps a custom Chromium build with source-level fingerprint changes and Playwright- and Puppeteer-compatible APIs. It is aimed at browser automation and web scraping where standard browser automation is frequently flagged by bot detection.

PythonBrowser AutomationWeb Scraping
Added May 27, 2026 View details
ai-website-cloner-template: Recreate Websites as Next.js Apps

ai-website-cloner-template: Recreate Websites as Next.js Apps

A template that equips AI coding agents to inspect a website and recreate it as a Next.js application. It suits developers rebuilding sites they own, recovering lost source code, or studying real-world interface implementations.

AIAI AgentsTypeScript
Added May 26, 2026 View details
firecrawl: Search, Scrape, and Crawl the Web for AI

firecrawl: Search, Scrape, and Crawl the Web for AI

Firecrawl turns web pages and sites into content that AI applications can use, including Markdown and structured data. It suits developers building agents or data pipelines that need web search, extraction, and crawling through an API or self-hosted deployment.

TypeScriptAIAI Agents
Added May 13, 2026 View details
Cheerio: Fast and Flexible HTML/XML Parsing and Manipulation Library

Cheerio: Fast and Flexible HTML/XML Parsing and Manipulation Library

Cheerio is a popular library for parsing and manipulating HTML and XML documents in Node.js. It provides a jQuery-like API, making it easy to select, traverse, and modify elements with proven syntax. Known for its blazingly fast performance and incredible flexibility, Cheerio is an excellent choice for web scraping and server-side DOM manipulation.

CheerioHTML ParserWeb Scraping
Added May 13, 2026 View details
Previous Page 1 Next

Related topics

OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️