Repository History

10 repositories tagged with web-scraping

Topic: web-scraping
proxy-list: A Daily Updated List of Free Proxy Servers

proxy-list: A Daily Updated List of Free Proxy Servers

The `clarketm/proxy-list` repository provides a comprehensive, daily updated collection of free, public, forward proxy servers. It offers various formats for easy access, including raw IP:PORT lists and detailed information on country, anonymity, and type. This resource is invaluable for developers and users needing reliable proxy access for various networking tasks.

Analyzed Jul 8, 2026
View Details
PinchTab: High-Performance Browser Automation for AI Agents

PinchTab: High-Performance Browser Automation for AI Agents

PinchTab is a high-performance browser automation bridge and multi-instance orchestrator, designed to give AI agents direct control over Chrome. Built in Go, it offers advanced stealth injection, real-time dashboards, and token-efficient web interaction. It supports both headless and headed modes, enabling robust and secure automation workflows for various applications.

Analyzed Jun 21, 2026
View Details
CloakBrowser: Stealth Chromium for Unblockable Web Scraping and Automation

CloakBrowser: Stealth Chromium for Unblockable Web Scraping and Automation

CloakBrowser is a powerful, open-source stealth Chromium browser engineered to bypass advanced bot detection systems. It achieves unparalleled stealth through C++ source-level fingerprint patches, making it appear as a normal browser and passing over 30 detection tests. Designed as a drop-in replacement for Playwright and Puppeteer, it simplifies web automation for AI agents, web scraping, and more.

Analyzed May 27, 2026
View Details
AI Website Cloner Template: Clone Websites with AI Coding Agents

AI Website Cloner Template: Clone Websites with AI Coding Agents

The AI Website Cloner Template is an innovative open-source project that leverages AI coding agents to reverse-engineer any website into a clean, modern Next.js codebase. It enables users to clone entire websites with a single command, extracting design tokens, assets, and reconstructing sections in parallel. This tool is ideal for platform migration, recovering lost source code, or learning web development by deconstructing live sites.

Analyzed May 26, 2026
View Details
Cheerio: Fast and Flexible HTML/XML Parsing and Manipulation Library

Cheerio: Fast and Flexible HTML/XML Parsing and Manipulation Library

Cheerio is a popular library for parsing and manipulating HTML and XML documents in Node.js. It provides a jQuery-like API, making it easy to select, traverse, and modify elements with proven syntax. Known for its blazingly fast performance and incredible flexibility, Cheerio is an excellent choice for web scraping and server-side DOM manipulation.

Analyzed May 13, 2026
View Details
Scraperr: A Powerful Self-Hosted Web Scraping Solution

Scraperr: A Powerful Self-Hosted Web Scraping Solution

Scraperr is a powerful self-hosted web scraping solution that allows users to extract data from websites without writing a single line of code. It features XPath-based extraction, queue management, domain spidering, and various data export options. This tool provides a comprehensive platform for efficient and controlled web data collection.

Analyzed Feb 16, 2026
View Details
brightdata-mcp: Empowering AI with Real-time Web Access and Data Scraping

brightdata-mcp: Empowering AI with Real-time Web Access and Data Scraping

The brightdata-mcp is a powerful Model Context Protocol (MCP) server developed by Bright Data, designed to give AI agents real-time web access. It provides an all-in-one solution for seamless public web interaction, ensuring Large Language Models (LLMs) can access live information without encountering blocks or CAPTCHAs. This open-source project offers robust web scraping, browser automation, and data extraction capabilities.

Analyzed Jan 22, 2026
View Details
13ft: Self-Hosted Paywall Bypass and Ad Blocker

13ft: Self-Hosted Paywall Bypass and Ad Blocker

13ft is a powerful, self-hosted Python application designed to bypass paywalls and block ads on various websites, including those that services like 12ft.io might miss. It operates by impersonating GoogleBot to access the full content of articles. This open-source tool offers a flexible solution for users seeking to read restricted content.

Analyzed Dec 10, 2025
View Details
Browserable: Open Source Browser Automation for AI Agents

Browserable: Open Source Browser Automation for AI Agents

Browserable is an open-source and self-hostable library designed to empower AI agents with advanced browser automation capabilities. It enables agents to navigate websites, fill out forms, click buttons, and extract information efficiently. With a strong performance on Web Voyager benchmarks, Browserable provides a robust foundation for building intelligent AI-driven web interactions.

Analyzed Oct 12, 2025
View Details
AnyCrawl: A High-Performance Node.js/TypeScript Web Crawler for LLM Data

AnyCrawl: A High-Performance Node.js/TypeScript Web Crawler for LLM Data

AnyCrawl is a powerful Node.js/TypeScript web crawler designed to transform websites into LLM-ready data. It excels at extracting structured SERP results from various search engines and features native multi-threading for efficient bulk processing, making it ideal for large-scale data collection.

Analyzed Oct 12, 2025
View Details
Previous Page 1 Next
OS
OSRepos

Analysis and discovery of open source repositories. Find interesting projects and follow their updates.

Monitor your website with YourWebsiteScore

OSRepos shares public repositories for knowledge and discovery only. Any installation, execution, configuration, or use of third-party repository code is at your own risk. Always review source code, dependencies, licenses, and security implications before running anything.

© 2025 OSRepos. Built with Nuxt 3 and lots of ❤️