Web Crawlers
Web crawlers are tools that automatically visit web pages, follow links, and collect information across a site or the wider internet. They help build search indexes, monitor changes, map site structure, and gather content for research or analysis. Crawling can range from fetching a small set of pages to coordinating large, distributed workloads, while respecting access rules and managing request rates is important for reliable use.
Open source options include libraries, command-line tools, and configurable crawling frameworks, with features for scheduling, link discovery, rendering, and data storage. When choosing one, consider supported languages, site complexity, performance needs, integration requirements, license, documentation, and maintenance activity. These tools are useful to developers, researchers, and organizations that need repeatable ways to discover and process web content.
2 repositories · updated April 7, 2026

OmniParse: Ingest, Parse, and Optimize Data for GenAI Frameworks
OmniParse is a powerful platform designed to ingest, parse, and optimize any unstructured data, from documents to multimedia, into structured, actionable formats. It enhances compatibility with GenAI frameworks, preparing data for applications like RAG and fine-tuning. This tool simplifies the complex process of data preparation for AI, making it accessible and efficient.

PinescriptV6-docs-crawler: Python Tool for Pine Script V6 Documentation
PinescriptV6-docs-crawler is a Python tool designed to crawl and process TradingView's Pine Script V6 documentation. Utilizing the Crawl4Ai framework, it efficiently extracts, cleans, and organizes this documentation into searchable markdown files. This makes it significantly easier for developers to reference and analyze Pine Script features and syntax.