proxy-compare.com

Blog · How it works

How to fix a 403 when scraping

Proxy Compare requested the front page and robots.txt of ten large sites on 28 Aug 2026 from a residential connection, then repeated every refusal with a python-urllib user agent. A 403 is several different failures wearing one status code, and which one you have decides what helps.

3 Sept 2026 · 8 min read

Point in time. The figures below were read on 28 Aug 2026 and are not updated after publication. For current numbers see the comparison table.

On this page (9)
How to fix a 403 when scraping

A 403 means the server understood your request and refused it. In a scraping context that refusal comes from one of two very different places, and they are indistinguishable from the status code alone.

Either the site itself turned you away because your client announced what it was, which is a one-header fix. Or a bot-management product sitting in front of the site made the decision, in which case it is reading your TLS handshake, your header order and your address, and no single header will change its mind.

The standard advice, set a user agent and buy proxies, fixes exactly one of those. Which is why the first step is not to change anything.

Find out which you have before changing anything. It takes one request and it is written in the response.

Step 1: read who refused you

Look at the Server header and the Set-Cookie line, not the status code.

When we asked ten large sites for their front page, two refused, and neither refusal came from the site itself:

text
indeed.com   403   Server: cloudflare     Set-Cookie: __cf_bm=...   CF-RAY: ...
                   page title: "security check - indeed.com"

ebay.com     403   Server: AkamaiGHost    Set-Cookie: bm_s=...
                   page title: "error page | ebay"

The cookie is the tell. __cf_bm is Cloudflare Bot Management, bm_s is Akamai Bot Manager. A 403 that arrives with one of those is a bot-management decision, and that tells you the class of problem before you touch a header.

A 403 with no such cookie, from the origin's own server, is usually the simpler case below.

Do not read the page copy for this. Here is what eBay actually returned:

The page eBay returned with its 403: a large photograph beside the heading SORRY and the line Something went wrong on our end, then a red monospace reference code and a Go to homepage button. Nothing on the page mentions automation, rate limits or bot detection.

It is dressed as a server fault. "Something went wrong on our end" is what a 500 says, and there is nothing on the page about automation, rate limits or bot detection. The red monospace string is an Akamai trace reference and is the only part their support will care about.

The status code and the cookie told the truth. The copy did not.

Step 2: if there is no bot manager, try the user agent

This is the advice everyone gives, and on the right failure it works completely.

Wikipedia refuses requests with its default user agent, which announces the library:

RequestResponse
requests, default user agent403, 126 characters
requests, Chrome 151 user agent200, 239,671 characters
python
headers = {"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
                         "AppleWebKit/537.36 (KHTML, like Gecko) "
                         "Chrome/151.0.0.0 Safari/537.36"}
r = requests.get(url, headers=headers)

One header, complete reversal.

Step 3: when a header will not help

We repeated both bot-managed refusals with a python-urllib user agent instead of Chrome:

SiteChrome UApython-urllib UA
indeed.com403403
ebay.com403403

Byte counts barely moved. A bot-management product is reading TLS fingerprint, header order, and browser signals, and one string is not the thing it is deciding on.

What changes the picture there is the whole client, not one field. The browser signals these products can read are the ones we measured across seven anti-detect tools, and a tool like curl_cffi with impersonate="chrome" exists to match a real browser's TLS handshake rather than only its user-agent string. It is also the cheapest thing to run, at an eighth of a browser's bandwidth.

Sampling more than once

We went back and repeated the identical request eight times against each, and the two sites turned out to be doing different things:

Site403200Refusal
ebay.com80deterministic
indeed.com53intermittent

eBay refused all eight, at about 1,830 bytes each. indeed.com refused five and served the real 600 KB page the other three, from the same address with the same headers, seconds apart.

Our own first pass got this wrong. One request made indeed.com look like a stable wall when it is closer to a weighted coin, and that difference decides your retry strategy: against eBay a retry is wasted bandwidth, against indeed.com it is most of the fix.

So sample before concluding, in both directions. A single request understates a rate limiter and overstates a probabilistic blocker.

Step 4: check whether you need a proxy at all

Uncomfortable for a proxy comparison site to say plainly: seven of these ten sites served an ordinary request from an ordinary connection, with no proxy.

If access is the only reason you want one, the check is one request. Where a proxy does change things is rate limiting and address reputation, which need volume to show up and which a single request cannot tell you about.

Where an address is genuinely the problem, what per-gigabyte pricing looks like across the market is on the residential proxy page with the date each figure was read. We rank on published price and measured behaviour, and no affiliate relationship enters a sort.

What robots.txt will and will not tell you

Not this. We checked, and the relationship is inverted:

Siterobots.txt disallows all for *Front page
linkedin.comyes200
x.comyes200
indeed.comno403 on 5 of 8
ebay.comno403 on 8 of 8
amazon.comno202

robots.txt is a policy statement from whoever owns the crawling relationship. A 403 is an engineering decision by a product in front of the origin. Different teams, different purposes, no obligation to agree.

That does not make the policy irrelevant. LinkedIn's file is explicit:

text
# Notice: The use of robots or other automated means to access LinkedIn without
# the express permission of LinkedIn is strictly prohibited.

It also publishes a route to ask, which almost nothing else does: whitelist-crawl@linkedin.com. Being served a 200 is not permission, and the terms of use are the document governing whether you may, not the status code.

The status code that is neither

Amazon never returned 200 to us at all:

User agentResponse
Chrome 151202, 2,007 bytes
python-urllib/3.14503, 2,671 bytes

A 202 is "accepted", and two kilobytes is not a front page. If you are branching on status == 403 you will miss this shape entirely, so branch on "did I get the content I expected" instead.

The check, as a script

Everything above is one request and three fields off the response. This prints the verdict for any URL:

python
import requests

UA = ("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
      "(KHTML, like Gecko) Chrome/151.0.0.0 Safari/537.36")

BOT_COOKIES = {
    "__cf_bm":   "Cloudflare Bot Management",
    "bm_s":      "Akamai Bot Manager",
    "bm_sv":     "Akamai Bot Manager",
    "datadome":  "DataDome",
    "_px":       "PerimeterX",
}

def diagnose(url):
    r = requests.get(url, headers={"User-Agent": UA}, timeout=20)
    server = r.headers.get("server", "")
    cookies = "; ".join(r.cookies.keys())

    hit = next((v for k, v in BOT_COOKIES.items()
                if any(c.startswith(k) for c in r.cookies.keys())), None)

    print(f"{url}\n  status {r.status_code}  server {server}  bytes {len(r.content):,}")
    if hit:
        print(f"  -> bot manager: {hit}. A user agent will not fix this.")
    elif r.status_code == 403:
        print("  -> refused with no bot-manager cookie. Try the user agent first.")
    elif r.status_code != 200:
        print(f"  -> neither served nor refused. Branch on content, not on 403.")
    else:
        print("  -> served. No access problem to solve from this address.")

for site in ["https://www.indeed.com/", "https://www.ebay.com/",
             "https://www.amazon.com/", "https://en.wikipedia.org/"]:
    diagnose(site)

Against those four sites it prints:

text
https://www.indeed.com/
  status 403  server cloudflare  bytes 27,645
  -> bot manager: Cloudflare Bot Management. A user agent will not fix this.
https://www.ebay.com/
  status 403  server AkamaiGHost  bytes 1,831
  -> bot manager: Akamai Bot Manager. A user agent will not fix this.
https://www.amazon.com/
  status 202  server CloudFront  bytes 0
  -> neither served nor refused. Branch on content, not on 403.
https://en.wikipedia.org/
  status 200  server mw-web.eqiad.main-c7c945899-5f6lw  bytes 261,384
  -> served. No access problem to solve from this address.

Two of those numbers differ from the table earlier: that one was measured with urllib and this run uses requests, which handles redirects and compression differently. Amazon returning 202 with a zero-length body here rather than 2,007 bytes is the same refusal wearing different clothes, and it is a good argument for branching on content rather than on a status code or a size.

The proxy behind it is the recurring cost, and per-gigabyte prices vary more between providers than these tools vary between each other. What each charges, with the date it was read, is on the provider table.

Limitations

One address, one day. The ten-site sweep was a single request each; only indeed.com and ebay.com were sampled eight times, and that sample is what uncovered the difference between them.

A small sample errs in both directions. It understates a rate limiter, because per-path rules and address reputation need volume to appear and we generated none. And it overstates a probabilistic blocker: one request had indeed.com filed as a wall when it refuses roughly three times in five. Eight requests is enough to separate deterministic from intermittent and not enough to characterise either.

We did not scrape any of these sites, fetched nothing beyond the front page and robots.txt, and make no claim about behaviour under load. We did not test whether a proxy clears the two bot-managed refusals, because doing that honestly needs many addresses across many providers.

Nothing here is legal advice about whether you may scrape a given site.

Sources

  1. LinkedIn robots.txt LinkedIn, read 28 Aug 2026
    # Notice: The use of robots or other automated means to access LinkedIn without the express permission of LinkedIn is strictly prohibited.

Questions

Why am I getting 403 when web scraping?

Two different causes that look identical. Either the site refused your client's default identity, which one header fixes, or a bot-management product refused the connection, which it does not. The response tells you which: a bot manager sets a cookie with the refusal.

Does changing the user agent fix a 403?

Sometimes completely and sometimes not at all. Wikipedia refused python-requests' default user agent and served the page immediately with a Chrome one. Two sites refusing us through Cloudflare and Akamai returned 403 to both, with the byte counts barely moving.

Does robots.txt tell me whether a site will block me?

Not in our reading, and the relationship was inverted. Of ten sites, the two that disallow everything for a generic crawler both served their front page, and the two that returned 403 both permit crawling. They are different documents written by different teams.

Do I need a proxy to fix a 403?

Not necessarily, and it is worth checking before buying. Seven of the ten sites we tried served a plain request from an ordinary connection. Where a bot manager is refusing you, an address change is one input among several it reads, so it is not automatic there either.

Does a 200 mean a site allows scraping?

No. It means one request from one address was served. LinkedIn and X both served us while stating in robots.txt that automated access is disallowed, and LinkedIn's file calls it strictly prohibited. Whether you may is a separate question from whether you can, and the terms of use govern it.