Scrape without getting blocked
You avoid blocks by behaving like a courteous guest: ask robots.txt where you may go, fetch no faster than one page a second, and slow down the moment the server says so. This page is the full procedure — permission, pace, backoff, and table extraction — with the working numbers that keep a crawl alive.
| Setting or limit | Value | Where it comes from |
|---|---|---|
| Polite default crawl rate | 1 request per second | Common crawler etiquette |
| Daily volume at that rate | 86,400 requests | 60 seconds × 60 minutes × 24 hours |
| Shared-host expectation | 1 request per 2–10 seconds | Hosting provider norms |
| Backoff after refusals | 1, 2, 4, 8, 16 s, capped at 30–60 s | Standard retry practice |
| robots.txt fetch limit | 500 KiB | RFC 9309 (September 2022) |
| Sitemap capacity | 50,000 URLs / 50 MB uncompressed | Sitemap protocol |
| Compression savings | 60–80% smaller HTML, CSS, JSON | gzip or brotli |
| Headless browser timeout | 30,000 ms default | Chromium navigation default |
Five ordered steps: check permission, plan from the sitemap, pace requests, fetch lean, and log every outcome.
The order matters. Scrapers get blocked when they fetch first and ask questions later, so permission and pacing come before parsing. Each step below names its typical failure, because knowing how a step breaks is what keeps the run alive.
Budget real time for the early steps. Reading robots.txt and mapping the sitemap takes under an hour for most sites and saves days of rework; skipping it is how jobs end with a banned address and no data.
Treat the log as a deliverable, not a byproduct. Status code, final URL, and fetch time per page are what let you spot a 429 spike or a soft block while it is still cheap to fix.
One request per second is the accepted polite default; shared hosts commonly expect one request every 2–10 seconds.
Speed is the variable that gets scrapers noticed. One request per second yields 86,400 requests over a full day, which covers most research jobs without ever placing you in a site's top-talkers list.
Context tightens the default. On shared hosting, where many unrelated sites sit behind one address, operators expect automated traffic to stay between one request per 2 seconds and one per 10 seconds — stretch your interval when the site's responses suggest strain.
Timeouts set the other boundary. Headless Chromium abandons a navigation after 30,000 milliseconds by default, and plain HTTP clients are usually set to 10–30 seconds; a fetch that hits its timeout counts as a refusal and joins the same backoff ladder as an explicit 429.
A 429 says slow down, not stop: honor Retry-After exactly, and double your wait when no timer is given.
HTTP 429 Too Many Requests was defined in RFC 6585 in 2012 to give servers a clean way to say slower. Unlike a 403, it is feedback with a timer attached, and crawlers that honor the timer are usually back to 200 responses within a minute.
The timer is Retry-After, delivered either as a plain number of seconds or as a full HTTP-date naming the moment to return. Parse both forms before your first run; a client that understands only integers will mishandle date-form responses.
When a refusal arrives without a header — a bare 429, a 503, a connection reset — use exponential backoff: 1, 2, 4, 8, 16 seconds, capped at 30–60. Six failed attempts cost at most about a minute of waiting; if refusals continue past that, park the job for hours, not seconds.
Two standards and one court case supply every hard limit cited on this page.
Six common failure patterns and the single correct response to each.
This guide relies only on primary material: RFC 9309, RFC 6585, the sitemap protocol limits, documented Chromium defaults, and the published hiQ Labs v. LinkedIn opinions.