Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
-
Updated
Sep 25, 2026 - Python
Open-source web crawler and scraper for LLMs and AI agents: any website into clean, LLM-ready Markdown. Run it yourself, or use Crawl4AI Cloud with one key.
Crawlee—A web scraping and browser automation library for Node.js to build reliable crawlers. In JavaScript and TypeScript. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Puppeteer, Playwright, Cheerio, JSDOM, and raw HTTP. Both headful and headless mode. With proxy rotation.
Crawlee—A web scraping and browser automation library for Python to build reliable crawlers. Extract data for AI, LLMs, RAG, or GPTs. Download HTML, PDF, JPG, PNG, and other files from websites. Works with Parsel, BeautifulSoup, Playwright, and raw HTTP. Both headful and headless mode. With proxy rotation.
Declarative data automation language and Go runtime for structured extraction workflows.
The All in One Framework to Build Undefeatable Scrapers
A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.
Best headless browser for AI agents. Lite, Fast, High-Compatibility. Built in Rust
Official repository for "Craw4LLM: Efficient Web Crawling for LLM Pretraining"
Stealth Chromium engine that stops scrapers and browser agents from getting blocked, with one line of code change.
A simple web scraper to extract Product Data and Pricing from Amazon
Library for Rapid (Web) Crawler and Scraper Development
Omnisci3nt is an open-source web reconnaissance and intelligence tool for extracting deep technical insights from domains, including subdomains, SSL certificates, exposed services, archived content, and configuration data. — Omnisci3nt gives you the full picture in seconds.
This is a Twitter Scraper which uses Selenium for scraping tweets. It is capable of scraping tweets from home, user profile, hashtag, query or search, and advanced searches.
Machine Learning Model for Sport Predictions (Football, Basketball, Baseball, Hockey, Soccer & Tennis)
A simple but powerful web crawler library for .NET
⚡ Ayakashi.io - The next generation web scraping framework
A tool for scraping emails, social media accounts, and much more information from websites using Google Search Results.
Apify SDK for Python—The official library for building Apify Actors: serverless cloud programs for web scraping, browser automation, data processing, and AI agents. Manages the Actor lifecycle, storages (datasets, key-value stores, request queues), events, proxies, and pay-per-event monetization. Built on top of the the Apify API Client.
Official Search1API MCP server for web search, news, crawling, sitemaps, and trends—hosted with OAuth 2.1 or local via npm.
Scrapy Training companion code
To associate your repository with the web-crawling topic, visit your repo's landing page and select "manage topics."