How to use a proxy in Scrapy
Proxy Compare ran Scrapy 2.18.0 on Python 3.14.7 on 28 Aug 2026 against a local proxy that logs every request and whether its credentials matched, and confirmed all three attachment methods below carry authentication on every request rather than only the first.
3 Sept 2026 · 5 min read
Point in time. The figures below were read on 28 Aug 2026 and are not updated after publication. For current numbers see the comparison table.

What is Scrapy, and why proxies work the way they do in it
Scrapy is a Python crawling framework rather than a browser. It fetches HTML over HTTP and parses it, with a scheduler, a retry system, concurrency limits and a middleware chain around every request. It runs no JavaScript, which makes it enormously cheaper than a headless browser and useless against a page that builds itself client-side.
That architecture explains something that confuses people arriving from
Playwright or Selenium: Scrapy has no proxy setting. There is no PROXY line
in settings.py and no constructor argument, which is why every guide to this
shows you a different thing and none of them explains the shape.
The shape is one key. A request carries meta["proxy"], and Scrapy's built-in
HttpProxyMiddleware reads it. Everything below is a different way of writing
that key.
Then there is a change in 2.18 that can skip whichever way you picked, and that is at the end.
Setting a proxy for one request
The direct form. Credentials go in the URL if the proxy needs them:
import scrapy
class MySpider(scrapy.Spider):
name = "example"
async def start(self):
yield scrapy.Request(
"http://example.com/",
meta={"proxy": "http://user:pass@gate.example.com:7000"},
)
def parse(self, response):
yield {"url": response.url, "status": response.status}
Note async def start, not start_requests. That matters more than it looks
and is the subject of the last section.
Setting one for every request
Writing meta on every Request gets old, and it misses retries and redirects
that Scrapy generates itself. A downloader middleware does not:
class ProxyMiddleware:
def process_request(self, request, spider):
request.meta["proxy"] = "http://user:pass@gate.example.com:7000"
# settings.py
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.ProxyMiddleware": 543,
}
process_request runs for every outgoing request, which is the reason to
prefer this over decorating each one by hand.
Rotating a pool
Same hook, one line different. Cycle a list instead of returning a constant:
import itertools
class RotatingProxyMiddleware:
def __init__(self):
self.pool = itertools.cycle([
"http://user:pass@gate1.example.com:7000",
"http://user:pass@gate2.example.com:7000",
"http://user:pass@gate3.example.com:7000",
])
def process_request(self, request, spider):
request.meta["proxy"] = next(self.pool)
We ran this over three URLs and all three reached our proxy authenticated.
Round-robin is the simplest policy and not always the right one. It spreads
load evenly and it will keep handing out an address that has started failing,
because nothing here reads the response. If you need a dead proxy to drop out of
the pool, that logic goes in process_response and process_exception, and it
is the part worth writing yourself rather than copying.
A note on what you are rotating. With most residential providers you are not given a list of gateways at all: you get one endpoint and the rotation happens on their side, keyed by something in the username. In that case the pool above is the wrong tool and the session token is the right one. What each provider does is on its provider page with the date we read it.
Authenticating
Two forms, and both worked in our test.
Credentials in the URL, which is the common one:
request.meta["proxy"] = "http://user:pass@gate.example.com:7000"
An explicit header, with the proxy URL left bare:
import base64
token = base64.b64encode(b"user:pass").decode()
yield scrapy.Request(
url,
meta={"proxy": "http://gate.example.com:7000"},
headers={"Proxy-Authorization": f"Basic {token}"},
)
Use the second when your password contains characters that fight the URL form, which is the same problem MoreLogin warns about explicitly in its own importer.
Scrapy sends the credentials on the first request rather than waiting for a 407 and repeating itself. That is worth knowing because browsers do the opposite: every one we metered spends a round trip being challenged first, which we measured in which tools actually send your proxy password.
Scrapy also parses http://user:pass@host:port correctly, which Chrome does not.
The identical string passed to Chrome's --proxy-server removes the proxy from
the request path entirely.
The 2.18 change that skips all of the above
Now the trap, because it makes every method on this page do nothing.
scrapy.Spider has no start_requests attribute at 2.18.0. Not deprecated,
absent:
>>> import scrapy
>>> hasattr(scrapy.Spider, "start_requests")
False
>>> hasattr(scrapy.Spider, "start")
True
So a start_requests method in your spider overrides nothing. It defines a
function that nothing calls, and no warning fires, because nothing was
deprecated. The replacement start is documented in the installed package as
"versionadded:: 2.13".
What that does to you depends on one unrelated line:
| Your spider | Pages crawled | Requests reaching the proxy |
|---|---|---|
start_requests, no start_urls | 0 | 0 |
start_requests + start_urls | 1 | 0 |
async def start | 1 | 1 |
The middle row is the one to worry about and it is the most common shape,
because start_urls is in every tutorial. The crawl succeeds. It returns
200, it returns real content, and the fetch came from your own address, because
start_urls is handled by Scrapy's default start behaviour which never saw the
meta you wrote.
That fails open. Compare Selenium, which given a credential it cannot parse refuses to load the page at all. A crawl that quietly stops using the proxy is the more expensive direction of the two.
The fix is two words: async def start.
Checking it actually worked
Do not read the spider log for this. In the middle row above the spider log reports a clean successful crawl.
Count requests at the proxy instead. If your provider has a usage dashboard, run a small job and see whether the number moved. If it did not move while your spider reported success, your requests went direct.
The proxy used here logs one line per request with whether the credentials matched, and it is in the repository with the spiders, so you can point a spider at it and see the truth in one run.
The proxy behind it is the recurring cost, and per-gigabyte prices vary more between providers than these tools vary between each other. What each charges, with the date it was read, is on the provider table.
If you are sizing a budget rather than choosing a tool, the cost calculator takes a request count and an average page weight and gives you the monthly figure.
Limitations
Scrapy 2.18.0 only. We did not establish in which release start_requests was
removed, only that it is gone in this one, and we are not guessing at a version
we did not check.
We did not test scrapy-rotating-proxies or other third-party middleware. A
middleware setting meta["proxy"] in process_request runs after the request
exists, so it is unaffected by the start change, but we measured our own
middleware rather than theirs and will not speak for code we did not run.
We did not test proxy performance, failure handling under load, or what happens when a pool member dies mid-crawl.
Sources
- Scrapy PyPI, read 28 Aug 2026
scrapy 2.18.0
Questions
How do I set a proxy in Scrapy?
Put it in a request's meta dictionary as meta['proxy'], including credentials if it needs them. Scrapy's built-in HttpProxyMiddleware reads that key and routes the request. There is no proxy setting in settings.py, which is why every approach below is a variation on writing that one key.
How do I rotate proxies in Scrapy?
A downloader middleware that sets request.meta['proxy'] in process_request. That runs for every request including retries and redirects, so a pool cycled there is applied everywhere without touching your spider. We verified this against three URLs and all three arrived at the proxy authenticated.
How do I use an authenticated proxy in Scrapy?
Two ways, and both worked in our test. Put the credentials in the meta URL as http://user:pass@host:port, or send an explicit Proxy-Authorization header with a base64 Basic token and leave the meta URL bare. Scrapy sends the credentials on the first request rather than waiting to be challenged.
Why is my Scrapy spider crawling 0 pages?
If your requests come from start_requests, that method is no longer called. At 2.18.0 the Spider class has no start_requests attribute, so defining one leaves a method nothing invokes. Rename it to an async start method.
Does Scrapy need a proxy at all?
Only if your target refuses you or rate-limits your address. We checked ten large sites and seven served a plain request from an ordinary connection. Check before buying, because access is the reason most people think they need one and it is often not the reason they do.
