requests is still the quickest way to fetch pages from Python, and it supports proxies out of the box. This guide covers the setup that works, the one detail that trips most people up (it's the https key), and the patterns you'll want once you scale past a few requests.
All examples use the Residential gateway geo.crawlproxies.com. Grab your username and password from the generator and swap them in.
The basic setup
Pass a proxies dict with both keys:
import requests
PROXY = "http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
proxies = {"http": PROXY, "https": PROXY}
r = requests.get("https://ipinfo.io/json", proxies=proxies, timeout=30)
print(r.status_code, r.json()["ip"], r.json()["country"])The #1 mistake: thehttpskey still uses anhttp://proxy URL. The key is the scheme of the website you're visiting; the value is how to reach the proxy. HTTPS sites are tunnelled through the proxy withCONNECT, so your traffic stays encrypted end to end.
Always set a timeout. Without one, a single slow connection can hang your script forever.
Target a country, state or city
Targeting goes in the username, so it's just a string change:
def proxy_for(country=None, state=None, city=None):
user = "USERNAME"
if country:
user += f"-country-{country}"
if state:
user += f"-state-{state}"
if city:
user += f"-city-{city}"
url = f"http://{user}:PASSWORD@geo.crawlproxies.com:8080"
return {"http": url, "https": url}
r = requests.get("https://ipinfo.io/json", proxies=proxy_for("us", state="newyork"), timeout=30)
print(r.json()["region"]) # New YorkCountry codes are two letters (us, de, gb); states and cities are lowercase with no spaces (newyork, losangeles). See the geo-targeting guide for the details.
Rotating vs sticky
With no session in the username, each new connection gets a new IP. That's what you want for crawling lots of independent pages.
There's a catch: a requests.Session keeps connections alive and reuses them, so several requests in a row can leave from the same IP. If you need a fresh IP every time, either call requests.get() (a new connection each call) or ask the server to close the connection:
for _ in range(3):
r = requests.get("https://ipinfo.io/ip", proxies=proxies, timeout=30,
headers={"Connection": "close"})
print(r.text.strip()) # three different IPsWhen you need the same IP across requests (a login, a cart, a multi-step form), add a session id and a duration in seconds:
import uuid
session_id = uuid.uuid4().hex[:8]
STICKY = f"http://USERNAME-country-us-session-{session_id}-time-1800:PASSWORD@geo.crawlproxies.com:8080"
s = requests.Session()
s.proxies = {"http": STICKY, "https": STICKY}
print(s.get("https://ipinfo.io/ip", timeout=30).text)
print(s.get("https://ipinfo.io/ip", timeout=30).text) # same IP for up to 30 minutesUse a different session id per account or per worker. Sticky vs rotating sessions explains when to use which.
SOCKS5
Install the SOCKS extra once:
pip install "requests[socks]"Then use the SOCKS5 port (1080 on the Residential gateway) with the socks5h scheme:
SOCKS = "socks5h://USERNAME:PASSWORD@geo.crawlproxies.com:1080"
r = requests.get("https://ipinfo.io/json", proxies={"http": SOCKS, "https": SOCKS}, timeout=30)The h in socks5h means DNS is resolved by the proxy too, so even your lookups don't leave from your own IP. Plain socks5:// resolves hostnames on your machine.
Retries with backoff
Residential IPs are real devices, so an occasional request will fail or a site will rate-limit you. Let urllib3 retry the transient cases with growing pauses:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
retry = Retry(
total=4,
backoff_factor=1, # 1s, 2s, 4s, 8s between attempts
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET", "HEAD"],
)
s = requests.Session()
s.mount("http://", HTTPAdapter(max_retries=retry))
s.mount("https://", HTTPAdapter(max_retries=retry))
s.proxies = proxies
r = s.get("https://example.com/", timeout=30)Only retry idempotent requests automatically. Retrying a POST that already went through can submit a form twice.
Many requests in parallel
requests is blocking, so use a thread pool to run requests side by side. Each thread gets its own connection, and with a rotating username each connection has its own IP:
from concurrent.futures import ThreadPoolExecutor
urls = [f"https://example.com/page/{i}" for i in range(1, 51)]
def fetch(url):
try:
r = requests.get(url, proxies=proxies, timeout=30)
return url, r.status_code
except requests.RequestException as e:
return url, f"error: {e.__class__.__name__}"
with ThreadPoolExecutor(max_workers=10) as pool:
for url, status in pool.map(fetch, urls):
print(status, url)Start with 5 to 10 workers and raise it slowly while watching your error rate. Hammering one site from many IPs at once is still hammering it.
Environment variables
requests also honours the standard proxy variables, which is handy for tools you don't want to edit:
export HTTP_PROXY="http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
export HTTPS_PROXY="http://USERNAME:PASSWORD@geo.crawlproxies.com:8080"
python your_script.pyOn Windows PowerShell use $env:HTTPS_PROXY = "http://..." instead.
Errors you might see
| Error | Usually means |
|---|---|
ProxyError ... 407 | Wrong username or password, or a typo in the targeting part |
ProxyError ... Unable to connect to proxy | Wrong host or port, or a firewall blocking outbound connections |
ConnectTimeout / ReadTimeout | A slow exit IP: retry it, and keep your timeouts sensible |
InvalidSchema: Missing dependencies for SOCKS support | Run pip install "requests[socks]" |
403 / 429 from the site | The site is blocking or rate-limiting: see our debugging checklist |
That's the whole toolkit. For browser automation in Python, Playwright takes the same credentials: see Playwright with authenticated proxies.



