sclawl / bot
About the crawler
DRAFT — not legal advice, verify with a qualified lawyer.
What it reads
sclawl fetches public HTML to compare a product's commitments with a provider's terms. A user starts each comparison. The crawler reads up to six HTML pages per site, plus discovery files such as robots.txt, sitemaps and security.txt. It does not execute page scripts, sign in, submit forms or read PDF content.
User agent: sclawl/0.1.0 (+https://sclawl.pages.dev/bot). Robots token: sclawl.
How to block the crawler
Add these lines to robots.txt on each origin you want to exclude:
User-agent: sclawl Disallow: /
Rules are checked before fetching pages and on redirect destinations. A missing robots.txt (404 or 410) is treated as no rules. Other failures stop the crawl on that origin. Public accessibility does not override a robots exclusion.
Questions
Contact: Pending confirmation. Please include the affected public URL. Do not send private documents or credentials.