Web Mine Tech — Web scraping and polite crawling | webminetech.com

Web Mine Tech — Web scraping and polite crawling | webminetech.com

How to Read a robots.txt File

robots.txt is the site's own list of off-limits paths; reading it correctly means finding your user-agent group and collecting every Disallow that applies.

Every robots.txt lives at the same address — the site root, fetched as /robots.txt — and since September 2022 its syntax has been formalized by RFC 9309. The file is plain text made of groups, each starting with one or more User-agent lines followed by rules.

Find the group that names your crawler first. If no group matches your agent's name, the wildcard group — User-agent: * — applies. Only one group wins: a specific match beats the wildcard, and rules from different groups never combine.

Inside the group, Disallow lines list path prefixes the agent must not fetch, and Allow lines carve exceptions back out. When rules overlap, the longest matching path wins, which is how an Allow can reopen a single folder inside a disallowed directory.

The signature mark of this site, drawn as a plate

On narrow screens, swipe or scroll the plate sideways.

Crawl-delay is not part of the RFC, but many sites still publish it and many crawlers honor it: it asks for a minimum number of seconds between requests. If a site sets crawl-delay to 10, that overrides your polite one-second default — the site has stated its price explicitly, and the crawl rate guide explains how to schedule around it.

Size matters: under RFC 9309 a crawler may refuse to fetch a robots.txt larger than 500 KiB. Huge files are edge cases, but if you cannot retrieve or parse the file at all, the safe reading is to proceed at minimum rate and maximum care, not to assume everything is open.

The Sitemap line points to the site's URL inventory. A single sitemap holds at most 50,000 URLs and 50 MB uncompressed; larger sites publish a sitemap index of child files. Enumerating sitemaps before crawling saves you from guessing URLs — and the status codes reference explains what to do with the answer each URL returns.

  • Fetch /robots.txt before any other page
  • Match your user-agent to a group; fall back to the wildcard
  • Collect every Disallow in that group
  • Apply Allow exceptions by longest path match
  • Honor crawl-delay if present
  • Follow the Sitemap line for the URL inventory

Further reading