Web Mine Tech — Web scraping and polite crawling | webminetech.com

Scrape without getting blocked

Web Mine Tech — Web scraping and polite crawling | webminetech.com

You avoid blocks by behaving like a courteous guest: ask robots.txt where you may go, fetch no faster than one page a second, and slow down the moment the server says so. This page is the full procedure — permission, pace, backoff, and table extraction — with the working numbers that keep a crawl alive.

robots.txtUser-agent groupCrawl-delayRetry-After
The signature mark of this site, drawn as a plate
1 req/spolite default crawl rate for one client
500 KiBrobots.txt size limit under RFC 9309
50,000URLs allowed in a single sitemap file

The working numbers

The working numbers behind every step of the procedure.
Setting or limitValueWhere it comes from
Polite default crawl rate1 request per secondCommon crawler etiquette
Daily volume at that rate86,400 requests60 seconds × 60 minutes × 24 hours
Shared-host expectation1 request per 2–10 secondsHosting provider norms
Backoff after refusals1, 2, 4, 8, 16 s, capped at 30–60 sStandard retry practice
robots.txt fetch limit500 KiBRFC 9309 (September 2022)
Sitemap capacity50,000 URLs / 50 MB uncompressedSitemap protocol
Compression savings60–80% smaller HTML, CSS, JSONgzip or brotli
Headless browser timeout30,000 ms defaultChromium navigation default
The working numbers behind every step of the procedure.

The procedure, start to finish

Five ordered steps: check permission, plan from the sitemap, pace requests, fetch lean, and log every outcome.

The order matters. Scrapers get blocked when they fetch first and ask questions later, so permission and pacing come before parsing. Each step below names its typical failure, because knowing how a step breaks is what keeps the run alive.

Budget real time for the early steps. Reading robots.txt and mapping the sitemap takes under an hour for most sites and saves days of rework; skipping it is how jobs end with a banned address and no data.

Treat the log as a deliverable, not a byproduct. Status code, final URL, and fetch time per page are what let you spot a 429 spike or a soft block while it is still cheap to fix.

  • Fetch /robots.txt and list every Disallow for your user-agent — about 10 minutes; a 403 here is a hard stop, not a puzzle
  • Open the sitemap named in robots.txt and enumerate target URLs — under an hour; files cap at 50,000 URLs and 50 MB uncompressed
  • Set a fixed 1-second delay between requests, or 2–10 seconds on shared hosting, before the first fetch
  • Enable gzip or brotli so a 1,000-page run moves 60–80% fewer bytes
  • Log status code, final URL after redirects, and fetch time for every page to catch 429 spikes and empty-200 soft blocks early

Pace: how fast is too fast

One request per second is the accepted polite default; shared hosts commonly expect one request every 2–10 seconds.

Speed is the variable that gets scrapers noticed. One request per second yields 86,400 requests over a full day, which covers most research jobs without ever placing you in a site's top-talkers list.

Context tightens the default. On shared hosting, where many unrelated sites sit behind one address, operators expect automated traffic to stay between one request per 2 seconds and one per 10 seconds — stretch your interval when the site's responses suggest strain.

Timeouts set the other boundary. Headless Chromium abandons a navigation after 30,000 milliseconds by default, and plain HTTP clients are usually set to 10–30 seconds; a fetch that hits its timeout counts as a refusal and joins the same backoff ladder as an explicit 429.

The signature mark of this site, drawn as a plate

When the server pushes back: 429 and 503

A 429 says slow down, not stop: honor Retry-After exactly, and double your wait when no timer is given.

HTTP 429 Too Many Requests was defined in RFC 6585 in 2012 to give servers a clean way to say slower. Unlike a 403, it is feedback with a timer attached, and crawlers that honor the timer are usually back to 200 responses within a minute.

The timer is Retry-After, delivered either as a plain number of seconds or as a full HTTP-date naming the moment to return. Parse both forms before your first run; a client that understands only integers will mishandle date-form responses.

When a refusal arrives without a header — a bare 429, a 503, a connection reset — use exponential backoff: 1, 2, 4, 8, 16 seconds, capped at 30–60. Six failed attempts cost at most about a minute of waiting; if refusals continue past that, park the job for hours, not seconds.

2012RFC 6585 defines HTTP 429 Too Many Requests and the Retry-After header.2019The US Ninth Circuit rules in hiQ Labs v. LinkedIn that scraping publicly reachable pages likelyfalls outside the CFAA.2022The Ninth Circuit reaffirms the hiQ holding after the case is sent back.September 2022RFC 9309 formalizes robots.txt syntax, including the 500 KiB fetch limit.
A six-attempt refusal sequence: waits of 1, 2, 4, 8, 16 seconds, then a long cool-down of hours.

Four dates that set the rules

Two standards and one court case supply every hard limit cited on this page.

2012RFC 6585 defines HTTP 429 Too Many Requests and the Retry-After header.2019The US Ninth Circuit rules in hiQ Labs v. LinkedIn that scraping publicly reachable pages likelyfalls outside the CFAA.2022The Ninth Circuit reaffirms the hiQ holding after the case is sent back.September 2022RFC 9309 formalizes robots.txt syntax, including the 500 KiB fetch limit.
Timeline: 4 dated entries

Refusals you will meet, by name

Six common failure patterns and the single correct response to each.

robots.txtUser-agent groupCrawl-delayRetry-After
The key terms of this guide, drawn to one scale

Primary sources behind this guide

This guide relies only on primary material: RFC 9309, RFC 6585, the sitemap protocol limits, documented Chromium defaults, and the published hiQ Labs v. LinkedIn opinions.