Outil IT95

TRAWL

I built TRAWL, a fast self-hosted scraper for apps and AI agents

r/SideProjectu/trawlarr30 septembre 2026

Capture du projet

Résumé

TRAWL est un outil de scraping open source et auto-hébergé qui permet de récupérer des données de pages web de manière rapide et fiable. Il utilise une combinaison de requêtes HTTP et de sessions de navigateur pour contourner les protections anti-scraping. TRAWL est conçu pour être utilisé avec des agents IA et des applications qui ont besoin de récupérer des données de pages web.

Pourquoi c’est intéressant

TRAWL est intéressant car il propose une solution rapide et fiable pour récupérer des données de pages web, ce qui peut être utile pour de nombreux cas d'utilisation. Il est également open source et auto-hébergé, ce qui signifie que les utilisateurs ont le contrôle total sur leurs données et peuvent personnaliser l'outil selon leurs besoins.

Comment Claude est utilisé

Le projet utilise MCP pour permettre aux agents IA de lire des pages web et d'extraire des informations de manière structurée.

Idées dérivées

  1. 01

    CAPTCHASolver

    Un outil de scraping spécialisé pour les sites web qui utilisent des CAPTCHAs pour empêcher le scraping.

  2. 02

    ScrapingCloud

    Un service de scraping en ligne qui utilise TRAWL pour fournir des données à des applications et des agents IA.

  3. 03

    DataExtractor

    Un plugin pour les navigateurs web qui utilise TRAWL pour extraire des informations de pages web et les présenter sous forme de données structurées.

Afficher le post original
You don't have to be building a "scraper" to run into scraping problems. Maybe your app needs to read a product page, your agent needs information from a site you gave it, or Prowlarr needs to reach an indexer. It works until the response is a JavaScript shell, a CAPTCHA, or a "Just a moment" page instead of the content you asked for. I've been building TRAWL (https://trawl.germondai.com) to handle that step. You give it a URL through an API or its MCP server. It tries a regular HTTP request first, reuses an existing browser session when possible, and moves to a fresh Camoufox browser when the site needs one. That keeps straightforward pages fast without giving up on pages that require more work. Part of the motivation was my own experience with the alternatives. FlareSolverr was often slow and failed on sites I needed. Byparr worked better for me, but requests could take even longer, and its Docker image was larger than I wanted to run. On the sites I use, TRAWL has been faster and more reliable, with a higher success rate. That's my experience rather than a claim about every site. TRAWL has flows for Cloudflare, Akamai and Imperva, plus support for challenges such as Turnstile, reCAPTCHA, hCaptcha and GeeTest. These are best-effort tools, of course. No scraper can promise that every site will work. The MCP tools let an AI client read a page as Markdown, extract structured records from a rendered page, take screenshots and inspect browser activity. There's also a local dashboard that shows request history, which tier was used and why a request failed. If you already use FlareSolverr with Prowlarr, TRAWL has a compatible endpoint. Speed was a big part of the design. In three selected same-machine comparisons (https://trawl.germondai.com/#compare), TRAWL responded faster than FlareSolverr and Byparr. On a Cloudflare + Turnstile demo, TRAWL took 5.9s, FlareSolverr took 13.2s, and Byparr took 18.2s. Results will vary with the site, IP and session state. It's open source and runs on your own machine: https://github.com/germondai/trawl I'd be interested to hear where your current scraping setup tends to fail, especially cases where a request looks successful but returns a challenge page.