Most scraper blocks are self-inflicted. A script fires 50 concurrent requests at the same domain, a 429 response appears, you add a time.sleep(1), and repeat. This is the wrong mental model.
Proper rate limiting isn’t just about avoiding blocks — it’s about being a good citizen, respecting server capacity, and building resilient scrapers that work indefinitely without manual intervention. This guide covers the full stack: adaptive delays, exponential backoff with jitter, token bucket control, robots.txt compliance, and per-domain concurrency limits.
Why Scrapers Get Blocked
Understanding the detection mechanisms helps you avoid them:
- Request velocity: Too many requests per second from a single IP
- Concurrency signatures: Dozens of simultaneous connections from the same IP
- Behavior patterns: No variation in timing, always hitting the same paths in order
- Missing headers: No
User-Agent, noAccept-Language, noReferer - Ignoring 429s: Continuing to request after explicit rate-limit signals
Fixing the velocity and concurrency problems eliminates most blocks. The rest is polish.
The Simplest Fix: Adaptive Delays
A fixed sleep(1) is predictable — and predictability is a fingerprint. Use a random delay within a range, and scale it based on response time:
import asyncio
import random
import time
import httpx
async def fetch_with_delay(
client: httpx.AsyncClient,
url: str,
min_delay: float = 1.0,
max_delay: float = 3.0,
) -> httpx.Response:
start = time.monotonic()
response = await client.get(url)
elapsed = time.monotonic() - start
# If the server is slow, it's under load — back off more
adaptive_base = max(min_delay, elapsed * 0.5)
delay = random.uniform(adaptive_base, adaptive_base + (max_delay - min_delay))
await asyncio.sleep(delay)
return response
The adaptive component means that when a server is responding slowly (high load), your scraper automatically backs off — polite behavior that also reduces your block risk.
Exponential Backoff with Jitter
When you hit a 429 or 503, don’t retry immediately. Use exponential backoff — double the wait time on each failure — with added jitter to prevent the “thundering herd” problem where multiple retrying scrapers all wait the exact same duration and flood the server simultaneously.
import asyncio
import random
import httpx
from typing import Optional
async def fetch_with_backoff(
client: httpx.AsyncClient,
url: str,
max_retries: int = 5,
base_delay: float = 1.0,
max_delay: float = 60.0,
) -> Optional[httpx.Response]:
for attempt in range(max_retries):
try:
response = await client.get(url, timeout=30.0)
if response.status_code == 429:
# Respect Retry-After header if present
retry_after = response.headers.get("Retry-After")
if retry_after:
wait = float(retry_after)
else:
wait = min(base_delay * (2 ** attempt), max_delay)
# Add jitter: ±25% of wait time
jitter = wait * 0.25 * random.uniform(-1, 1)
actual_wait = max(1.0, wait + jitter)
print(f"Rate limited. Waiting {actual_wait:.1f}s before retry {attempt + 1}/{max_retries}")
await asyncio.sleep(actual_wait)
continue
if response.status_code in (500, 502, 503, 504):
wait = min(base_delay * (2 ** attempt), max_delay)
jitter = wait * 0.25 * random.uniform(-1, 1)
await asyncio.sleep(max(1.0, wait + jitter))
continue
return response
except (httpx.ConnectTimeout, httpx.ReadTimeout):
if attempt == max_retries - 1:
raise
wait = min(base_delay * (2 ** attempt), max_delay)
await asyncio.sleep(wait)
return None # All retries exhausted
Retry-After: When a server sends a Retry-After header, it tells you exactly how long to wait. Always respect it — it’s the server telling you the fastest safe retry time, not a suggestion.
Token Bucket Rate Limiter
For sustained scraping across many URLs, you need a rate limiter that enforces a steady request rate — not just delays between individual requests. The token bucket algorithm is the standard approach: tokens accumulate at a fixed rate, each request consumes one token, requests wait if the bucket is empty.
import asyncio
import time
class TokenBucket:
"""
Allows `rate` requests per second sustained, with burst capacity up to `capacity`.
"""
def __init__(self, rate: float, capacity: float):
self.rate = rate # tokens per second
self.capacity = capacity # max tokens (burst ceiling)
self._tokens = capacity
self._last_refill = time.monotonic()
self._lock = asyncio.Lock()
async def acquire(self, tokens: float = 1.0) -> None:
async with self._lock:
await self._wait_for_tokens(tokens)
async def _wait_for_tokens(self, tokens: float) -> None:
while True:
self._refill()
if self._tokens >= tokens:
self._tokens -= tokens
return
# Calculate how long until enough tokens accumulate
deficit = tokens - self._tokens
wait_time = deficit / self.rate
await asyncio.sleep(wait_time)
def _refill(self) -> None:
now = time.monotonic()
elapsed = now - self._last_refill
self._tokens = min(self.capacity, self._tokens + elapsed * self.rate)
self._last_refill = now
# Usage: 2 requests/second, burst up to 5
limiter = TokenBucket(rate=2.0, capacity=5.0)
async def fetch_url(client: httpx.AsyncClient, url: str) -> httpx.Response:
await limiter.acquire()
return await client.get(url)
With a rate of 2 and capacity of 5, you can fire an initial burst of up to 5 requests, then sustain 2 per second indefinitely. Adjust these to match what the target site can handle.
Per-Domain Concurrency Limits
When scraping multiple domains concurrently, you need to limit concurrency per domain, not globally. A global semaphore of 10 might put all 10 slots on the same domain at once. Per-domain limits prevent this:
import asyncio
from collections import defaultdict
from urllib.parse import urlparse
class PerDomainLimiter:
def __init__(self, max_per_domain: int = 2):
self.max_per_domain = max_per_domain
self._semaphores: dict[str, asyncio.Semaphore] = defaultdict(
lambda: asyncio.Semaphore(max_per_domain)
)
def get_semaphore(self, url: str) -> asyncio.Semaphore:
domain = urlparse(url).netloc
return self._semaphores[domain]
domain_limiter = PerDomainLimiter(max_per_domain=2)
async def fetch_with_domain_limit(
client: httpx.AsyncClient,
url: str,
) -> httpx.Response:
semaphore = domain_limiter.get_semaphore(url)
async with semaphore:
return await client.get(url)
Respecting robots.txt
Robots.txt specifies what scrapers are allowed to access and often includes a Crawl-delay directive. Parsing and honoring it is both polite and legally safer in many jurisdictions:
import urllib.robotparser
from urllib.parse import urlparse
from functools import lru_cache
@lru_cache(maxsize=100)
def get_robots_parser(base_url: str) -> urllib.robotparser.RobotFileParser:
rp = urllib.robotparser.RobotFileParser()
rp.set_url(f"{base_url}/robots.txt")
rp.read()
return rp
def can_fetch(url: str, user_agent: str = "*") -> bool:
parsed = urlparse(url)
base_url = f"{parsed.scheme}://{parsed.netloc}"
rp = get_robots_parser(base_url)
return rp.can_fetch(user_agent, url)
def get_crawl_delay(url: str, user_agent: str = "*") -> float:
parsed = urlparse(url)
base_url = f"{parsed.scheme}://{parsed.netloc}"
rp = get_robots_parser(base_url)
delay = rp.crawl_delay(user_agent)
return delay if delay is not None else 1.0 # Default 1s if not specified
# Before scraping
url = "https://example.com/products/page-1"
if not can_fetch(url):
print(f"robots.txt disallows: {url}")
else:
delay = get_crawl_delay(url)
await asyncio.sleep(delay)
response = await client.get(url)
Caching robots.txt: The @lru_cache ensures you fetch robots.txt once per domain, not once per URL. This matters when scraping thousands of pages from the same domain.
Putting It All Together
A production scraper combines all these layers: robots.txt compliance, per-domain concurrency, token bucket rate limiting, and exponential backoff on failures:
import asyncio
import random
import time
import httpx
from urllib.parse import urlparse
class PoliteScraper:
def __init__(
self,
requests_per_second: float = 1.0,
max_concurrency_per_domain: int = 2,
user_agent: str = "PoliteBot/1.0 (+https://yoursite.com/bot)",
):
self.user_agent = user_agent
self.rate_limiter = TokenBucket(rate=requests_per_second, capacity=requests_per_second * 3)
self.domain_limiter = PerDomainLimiter(max_per_domain=max_concurrency_per_domain)
self._client: httpx.AsyncClient | None = None
async def __aenter__(self):
self._client = httpx.AsyncClient(
headers={
"User-Agent": self.user_agent,
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.5",
},
follow_redirects=True,
)
return self
async def __aexit__(self, *_):
if self._client:
await self._client.aclose()
async def get(self, url: str) -> httpx.Response | None:
if not can_fetch(url, self.user_agent):
print(f"Skipping (robots.txt): {url}")
return None
# Respect Crawl-delay from robots.txt
crawl_delay = get_crawl_delay(url, self.user_agent)
await self.rate_limiter.acquire()
semaphore = self.domain_limiter.get_semaphore(url)
async with semaphore:
response = await fetch_with_backoff(self._client, url)
# Honor crawl delay on top of rate limiting
await asyncio.sleep(crawl_delay + random.uniform(0, 0.5))
return response
# Usage
async def scrape_urls(urls: list[str]) -> list[str]:
results = []
async with PoliteScraper(requests_per_second=1.0) as scraper:
tasks = [scraper.get(url) for url in urls]
responses = await asyncio.gather(*tasks, return_exceptions=True)
for response in responses:
if isinstance(response, httpx.Response) and response.status_code == 200:
results.append(response.text)
return results
Using UnWeb to Skip the Complexity
For scraping that needs rate limiting and clean content extraction, UnWeb handles the rendering, rate limiting, and content normalization on its end. Your scraper becomes a simple async call:
import asyncio
import httpx
UNWEB_API_KEY = "your_api_key"
async def fetch_markdown(url: str, client: httpx.AsyncClient) -> str:
response = await client.get(
"https://api.unweb.info/v1/convert",
params={"url": url},
headers={"Authorization": f"Bearer {UNWEB_API_KEY}"},
timeout=30.0,
)
response.raise_for_status()
return response.json()["markdown"]
# UnWeb handles JS rendering, rate limiting, and returns clean Markdown
async def scrape_product_pages(urls: list[str]) -> list[str]:
async with httpx.AsyncClient() as client:
# Still be polite to UnWeb's API with a semaphore
semaphore = asyncio.Semaphore(5)
async def fetch_one(url):
async with semaphore:
return await fetch_markdown(url, client)
return await asyncio.gather(*[fetch_one(url) for url in urls])
The clean Markdown output also makes downstream parsing (tables, structured data, LLM extraction) significantly easier than working with raw HTML.
Clean content from any URL, no rate-limit headaches
UnWeb converts any web page — including JS-rendered content — to clean Markdown. Free tier available.
Quick Reference: Rate Limiting Checklist
- Use random delays in a range, not fixed sleeps
- Scale delays adaptively when servers are responding slowly
- Implement exponential backoff with jitter for 429 and 5xx responses
- Respect
Retry-Afterheaders — they tell you the safe retry time - Use a token bucket for sustained rate control, not just per-request delays
- Limit concurrency per domain, not just globally
- Parse robots.txt and honor
Crawl-delaydirectives - Set a descriptive
User-Agentwith contact info — it identifies you as a legitimate bot