All posts
Tutorials5 min readOct 11, 2026

Build an async web crawler in Python with httpx and rotating proxies

A complete, polite crawler in about 80 lines of Python: asyncio workers, a shared queue, robots.txt, retries with backoff and a new residential IP per request. Copy it, point it at a site, and grow it from there.

By crawlproxies

A crawler starts from one page, collects the links on it, visits those, and keeps going until it has seen what it needs. With asyncio and httpx, a handful of workers can fetch many pages at once while the proxy gives every request its own IP. This guide builds a complete crawler step by step; the full script is at the end.

You need Python 3.10+ and httpx:

bash
pip install httpx

Your proxy username and password are in the generator. The examples use the Residential gateway geo.crawlproxies.com:8080 and crawl books.toscrape.com, a sandbox site built for scraping practice.

The proxy client

httpx takes the proxy as one URL, credentials included:

python
import httpx

PROXY = "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"

limits = httpx.Limits(max_connections=8, max_keepalive_connections=0)
client = httpx.AsyncClient(proxy=PROXY, limits=limits, timeout=30, follow_redirects=True)

max_keepalive_connections=0 is what makes the IPs rotate. With a rotating username, every new connection gets a new IP, and httpx normally keeps connections open for the next request. With keep-alive off, each request opens its own connection and leaves from its own address. max_connections caps how many requests run at once.

Using an older httpx? Versions before 0.26 call the argument proxies= instead of proxy=.

The standard library's HTMLParser is enough to pull hrefs out of a page:

python
from html.parser import HTMLParser

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            href = dict(attrs).get("href")
            if href:
                self.links.append(href)

Links are relative more often than not, so join each one with the page's URL and drop the #fragment part, or the same page gets crawled twice:

python
from urllib.parse import urldefrag, urljoin

link = urldefrag(urljoin(page_url, href)).url

Retries with backoff

Some requests fail: a slow exit, a 429, a 503. Retrying after a short, growing pause fixes most of them, and because keep-alive is off, every retry also comes from a new IP:

python
import asyncio
import random

RETRY = {429, 500, 502, 503, 504}

async def fetch(client, url):
    for attempt in range(4):
        try:
            r = await client.get(url)
            if r.status_code not in RETRY:
                return r
        except httpx.TransportError:
            pass
        await asyncio.sleep(2 ** attempt + random.random())
    return None

Workers and the queue

The crawl itself is a queue of URLs and a few workers taking from it. Each worker fetches a page, records it and queues every new link on the same site:

python
async def worker(client, queue, seen, robots, pages):
    while True:
        url = await queue.get()
        try:
            if len(pages) >= MAX_PAGES or not robots.can_fetch("*", url):
                continue
            r = await fetch(client, url)
            if r is None or "text/html" not in r.headers.get("content-type", ""):
                continue
            pages.append((url, r.status_code))
            parser = LinkParser()
            parser.feed(r.text)
            for href in parser.links:
                link = urldefrag(urljoin(url, href)).url
                if urlparse(link).netloc == urlparse(START).netloc and link not in seen:
                    seen.add(link)
                    queue.put_nowait(link)
        finally:
            queue.task_done()

queue.task_done() in finally matters: the main function waits on queue.join(), which only returns once every queued URL has been marked done, skipped ones included.

robots.txt

Being a polite crawler means honouring the site's robots.txt. urllib.robotparser reads it; fetch it through the same client so it uses the proxy too:

python
from urllib.robotparser import RobotFileParser

robots = RobotFileParser()
r = await client.get(urljoin(START, "/robots.txt"))
robots.parse(r.text.splitlines() if r.status_code == 200 else [])

An empty rule set (no robots.txt) allows everything.

The full crawler

python
import asyncio
import random
from html.parser import HTMLParser
from urllib.parse import urldefrag, urljoin, urlparse
from urllib.robotparser import RobotFileParser

import httpx

PROXY = "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
START = "https://books.toscrape.com/"
MAX_PAGES = 200
WORKERS = 8
RETRY = {429, 500, 502, 503, 504}


class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            href = dict(attrs).get("href")
            if href:
                self.links.append(href)


async def fetch(client, url):
    for attempt in range(4):
        try:
            r = await client.get(url)
            if r.status_code not in RETRY:
                return r
        except httpx.TransportError:
            pass
        await asyncio.sleep(2 ** attempt + random.random())
    return None


async def worker(client, queue, seen, robots, pages):
    while True:
        url = await queue.get()
        try:
            if len(pages) >= MAX_PAGES or not robots.can_fetch("*", url):
                continue
            r = await fetch(client, url)
            if r is None or "text/html" not in r.headers.get("content-type", ""):
                continue
            pages.append((url, r.status_code))
            parser = LinkParser()
            parser.feed(r.text)
            for href in parser.links:
                link = urldefrag(urljoin(url, href)).url
                if urlparse(link).netloc == urlparse(START).netloc and link not in seen:
                    seen.add(link)
                    queue.put_nowait(link)
        finally:
            queue.task_done()


async def main():
    limits = httpx.Limits(max_connections=WORKERS, max_keepalive_connections=0)
    headers = {"User-Agent": "Mozilla/5.0 (compatible; my-crawler/1.0)"}
    async with httpx.AsyncClient(proxy=PROXY, limits=limits, timeout=30,
                                 follow_redirects=True, headers=headers) as client:
        robots = RobotFileParser()
        r = await client.get(urljoin(START, "/robots.txt"))
        robots.parse(r.text.splitlines() if r.status_code == 200 else [])

        queue = asyncio.Queue()
        seen, pages = {START}, []
        queue.put_nowait(START)
        tasks = [asyncio.create_task(worker(client, queue, seen, robots, pages)) for _ in range(WORKERS)]
        await queue.join()
        for t in tasks:
            t.cancel()

    print(f"crawled {len(pages)} pages")
    for url, status in pages[:10]:
        print(status, url)


asyncio.run(main())

Run it with python crawler.py. Before pointing it at another site, read that site's terms, keep WORKERS low and add a delay if responses slow down.

Where to take it next

  • Save what you find. Parse the fields you need from r.text (with selectolax, lxml or BeautifulSoup) and write them out as you go, instead of keeping pages in memory.
  • Target a country. Add -country-us (or any code) to the username so the site sees local visitors; see the geo-targeting guide.
  • Pages behind a login. Switch those requests to a sticky session so they keep one IP: sticky vs rotating sessions.
  • JavaScript-heavy pages. Render them with a headless browser: Playwright.
  • Blocks and CAPTCHAs. Work through the debugging checklist.

Prefer a framework? Scrapy does the queue, retries and throttling for you.

Written by
crawlproxies
Create account