Web Scraping Rate Limiting in Python | UnWeb

Most scraper blocks are self-inflicted. A script fires 50 concurrent requests at the same domain, a 429 response appears, you add a time.sleep(1), and repeat. This is the wrong mental model.

Proper rate limiting isn’t just about avoiding blocks — it’s about being a good citizen, respecting server capacity, and building resilient scrapers that work indefinitely without manual intervention. This guide covers the full stack: adaptive delays, exponential backoff with jitter, token bucket control, robots.txt compliance, and per-domain concurrency limits.

Why Scrapers Get Blocked

Understanding the detection mechanisms helps you avoid them:

Fixing the velocity and concurrency problems eliminates most blocks. The rest is polish.

The Simplest Fix: Adaptive Delays

A fixed sleep(1) is predictable — and predictability is a fingerprint. Use a random delay within a range, and scale it based on response time:

import asyncio
import random
import time
import httpx

async def fetch_with_delay(
    client: httpx.AsyncClient,
    url: str,
    min_delay: float = 1.0,
    max_delay: float = 3.0,
) -> httpx.Response:
    start = time.monotonic()
    response = await client.get(url)
    elapsed = time.monotonic() - start

    # If the server is slow, it's under load — back off more
    adaptive_base = max(min_delay, elapsed * 0.5)
    delay = random.uniform(adaptive_base, adaptive_base + (max_delay - min_delay))
    await asyncio.sleep(delay)

    return response

The adaptive component means that when a server is responding slowly (high load), your scraper automatically backs off — polite behavior that also reduces your block risk.

Exponential Backoff with Jitter

When you hit a 429 or 503, don’t retry immediately. Use exponential backoff — double the wait time on each failure — with added jitter to prevent the “thundering herd” problem where multiple retrying scrapers all wait the exact same duration and flood the server simultaneously.

import asyncio
import random
import httpx
from typing import Optional

async def fetch_with_backoff(
    client: httpx.AsyncClient,
    url: str,
    max_retries: int = 5,
    base_delay: float = 1.0,
    max_delay: float = 60.0,
) -> Optional[httpx.Response]:
    for attempt in range(max_retries):
        try:
            response = await client.get(url, timeout=30.0)

            if response.status_code == 429:
                # Respect Retry-After header if present
                retry_after = response.headers.get("Retry-After")
                if retry_after:
                    wait = float(retry_after)
                else:
                    wait = min(base_delay * (2 ** attempt), max_delay)
                # Add jitter: ±25% of wait time
                jitter = wait * 0.25 * random.uniform(-1, 1)
                actual_wait = max(1.0, wait + jitter)
                print(f"Rate limited. Waiting {actual_wait:.1f}s before retry {attempt + 1}/{max_retries}")
                await asyncio.sleep(actual_wait)
                continue

            if response.status_code in (500, 502, 503, 504):
                wait = min(base_delay * (2 ** attempt), max_delay)
                jitter = wait * 0.25 * random.uniform(-1, 1)
                await asyncio.sleep(max(1.0, wait + jitter))
                continue

            return response

        except (httpx.ConnectTimeout, httpx.ReadTimeout):
            if attempt == max_retries - 1:
                raise
            wait = min(base_delay * (2 ** attempt), max_delay)
            await asyncio.sleep(wait)

    return None  # All retries exhausted

Retry-After: When a server sends a Retry-After header, it tells you exactly how long to wait. Always respect it — it’s the server telling you the fastest safe retry time, not a suggestion.

Token Bucket Rate Limiter

For sustained scraping across many URLs, you need a rate limiter that enforces a steady request rate — not just delays between individual requests. The token bucket algorithm is the standard approach: tokens accumulate at a fixed rate, each request consumes one token, requests wait if the bucket is empty.

import asyncio
import time

class TokenBucket:
    """
    Allows `rate` requests per second sustained, with burst capacity up to `capacity`.
    """
    def __init__(self, rate: float, capacity: float):
        self.rate = rate          # tokens per second
        self.capacity = capacity  # max tokens (burst ceiling)
        self._tokens = capacity
        self._last_refill = time.monotonic()
        self._lock = asyncio.Lock()

    async def acquire(self, tokens: float = 1.0) -> None:
        async with self._lock:
            await self._wait_for_tokens(tokens)

    async def _wait_for_tokens(self, tokens: float) -> None:
        while True:
            self._refill()
            if self._tokens >= tokens:
                self._tokens -= tokens
                return
            # Calculate how long until enough tokens accumulate
            deficit = tokens - self._tokens
            wait_time = deficit / self.rate
            await asyncio.sleep(wait_time)

    def _refill(self) -> None:
        now = time.monotonic()
        elapsed = now - self._last_refill
        self._tokens = min(self.capacity, self._tokens + elapsed * self.rate)
        self._last_refill = now


# Usage: 2 requests/second, burst up to 5
limiter = TokenBucket(rate=2.0, capacity=5.0)

async def fetch_url(client: httpx.AsyncClient, url: str) -> httpx.Response:
    await limiter.acquire()
    return await client.get(url)

With a rate of 2 and capacity of 5, you can fire an initial burst of up to 5 requests, then sustain 2 per second indefinitely. Adjust these to match what the target site can handle.

Per-Domain Concurrency Limits

When scraping multiple domains concurrently, you need to limit concurrency per domain, not globally. A global semaphore of 10 might put all 10 slots on the same domain at once. Per-domain limits prevent this:

import asyncio
from collections import defaultdict
from urllib.parse import urlparse

class PerDomainLimiter:
    def __init__(self, max_per_domain: int = 2):
        self.max_per_domain = max_per_domain
        self._semaphores: dict[str, asyncio.Semaphore] = defaultdict(
            lambda: asyncio.Semaphore(max_per_domain)
        )

    def get_semaphore(self, url: str) -> asyncio.Semaphore:
        domain = urlparse(url).netloc
        return self._semaphores[domain]


domain_limiter = PerDomainLimiter(max_per_domain=2)

async def fetch_with_domain_limit(
    client: httpx.AsyncClient,
    url: str,
) -> httpx.Response:
    semaphore = domain_limiter.get_semaphore(url)
    async with semaphore:
        return await client.get(url)

Respecting robots.txt

Robots.txt specifies what scrapers are allowed to access and often includes a Crawl-delay directive. Parsing and honoring it is both polite and legally safer in many jurisdictions:

import urllib.robotparser
from urllib.parse import urlparse
from functools import lru_cache

@lru_cache(maxsize=100)
def get_robots_parser(base_url: str) -> urllib.robotparser.RobotFileParser:
    rp = urllib.robotparser.RobotFileParser()
    rp.set_url(f"{base_url}/robots.txt")
    rp.read()
    return rp

def can_fetch(url: str, user_agent: str = "*") -> bool:
    parsed = urlparse(url)
    base_url = f"{parsed.scheme}://{parsed.netloc}"
    rp = get_robots_parser(base_url)
    return rp.can_fetch(user_agent, url)

def get_crawl_delay(url: str, user_agent: str = "*") -> float:
    parsed = urlparse(url)
    base_url = f"{parsed.scheme}://{parsed.netloc}"
    rp = get_robots_parser(base_url)
    delay = rp.crawl_delay(user_agent)
    return delay if delay is not None else 1.0  # Default 1s if not specified

# Before scraping
url = "https://example.com/products/page-1"
if not can_fetch(url):
    print(f"robots.txt disallows: {url}")
else:
    delay = get_crawl_delay(url)
    await asyncio.sleep(delay)
    response = await client.get(url)

Caching robots.txt: The @lru_cache ensures you fetch robots.txt once per domain, not once per URL. This matters when scraping thousands of pages from the same domain.

Putting It All Together

A production scraper combines all these layers: robots.txt compliance, per-domain concurrency, token bucket rate limiting, and exponential backoff on failures:

import asyncio
import random
import time
import httpx
from urllib.parse import urlparse

class PoliteScraper:
    def __init__(
        self,
        requests_per_second: float = 1.0,
        max_concurrency_per_domain: int = 2,
        user_agent: str = "PoliteBot/1.0 (+https://yoursite.com/bot)",
    ):
        self.user_agent = user_agent
        self.rate_limiter = TokenBucket(rate=requests_per_second, capacity=requests_per_second * 3)
        self.domain_limiter = PerDomainLimiter(max_per_domain=max_concurrency_per_domain)
        self._client: httpx.AsyncClient | None = None

    async def __aenter__(self):
        self._client = httpx.AsyncClient(
            headers={
                "User-Agent": self.user_agent,
                "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
                "Accept-Language": "en-US,en;q=0.5",
            },
            follow_redirects=True,
        )
        return self

    async def __aexit__(self, *_):
        if self._client:
            await self._client.aclose()

    async def get(self, url: str) -> httpx.Response | None:
        if not can_fetch(url, self.user_agent):
            print(f"Skipping (robots.txt): {url}")
            return None

        # Respect Crawl-delay from robots.txt
        crawl_delay = get_crawl_delay(url, self.user_agent)

        await self.rate_limiter.acquire()

        semaphore = self.domain_limiter.get_semaphore(url)
        async with semaphore:
            response = await fetch_with_backoff(self._client, url)
            # Honor crawl delay on top of rate limiting
            await asyncio.sleep(crawl_delay + random.uniform(0, 0.5))
            return response


# Usage
async def scrape_urls(urls: list[str]) -> list[str]:
    results = []
    async with PoliteScraper(requests_per_second=1.0) as scraper:
        tasks = [scraper.get(url) for url in urls]
        responses = await asyncio.gather(*tasks, return_exceptions=True)
        for response in responses:
            if isinstance(response, httpx.Response) and response.status_code == 200:
                results.append(response.text)
    return results

Using UnWeb to Skip the Complexity

For scraping that needs rate limiting and clean content extraction, UnWeb handles the rendering, rate limiting, and content normalization on its end. Your scraper becomes a simple async call:

import asyncio
import httpx

UNWEB_API_KEY = "your_api_key"

async def fetch_markdown(url: str, client: httpx.AsyncClient) -> str:
    response = await client.get(
        "https://api.unweb.info/v1/convert",
        params={"url": url},
        headers={"Authorization": f"Bearer {UNWEB_API_KEY}"},
        timeout=30.0,
    )
    response.raise_for_status()
    return response.json()["markdown"]

# UnWeb handles JS rendering, rate limiting, and returns clean Markdown
async def scrape_product_pages(urls: list[str]) -> list[str]:
    async with httpx.AsyncClient() as client:
        # Still be polite to UnWeb's API with a semaphore
        semaphore = asyncio.Semaphore(5)
        async def fetch_one(url):
            async with semaphore:
                return await fetch_markdown(url, client)
        return await asyncio.gather(*[fetch_one(url) for url in urls])

The clean Markdown output also makes downstream parsing (tables, structured data, LLM extraction) significantly easier than working with raw HTML.

Clean content from any URL, no rate-limit headaches

UnWeb converts any web page — including JS-rendered content — to clean Markdown. Free tier available.

Get your API key

Quick Reference: Rate Limiting Checklist