How to cut your scraping bandwidth bill
Proxy Compare wrote a proxy that counts bytes per connection and pulled the same article through six clients on 28 Aug 2026: curl_cffi, Playwright, Patchright, Camoufox, SeleniumBase and requests. The four changes below are ordered by what each one actually saved, largest first, measured rather than estimated.
3 Sept 2026 · 5 min read
Point in time. The figures below were read on 28 Aug 2026 and are not updated after publication. For current numbers see the comparison table.

Why bytes are the number that matters
Residential proxies are priced per gigabyte. Not per request, not per hour: per gigabyte of traffic that passes through them, which means your bill is decided by how much your scraper downloads rather than by how many pages it visits.
That makes it unlike almost every other cost in a scraping stack, and it has an awkward consequence. Two tools fetching the same page can differ eightfold in what they pull down, because a browser also fetches the images, fonts, stylesheets and scripts that make the page render, and an HTTP client fetches the document and stops. Both give you the same HTML to parse. Only one of them charges you for the pictures.
Almost every tool comparison measures speed and detection, and almost none measures the number that arrives on the invoice. So we metered it: one local proxy that counts bytes, one page, five tools. Below are four changes that move the number, largest saving first.
Here are four changes that move it, measured on one article through a proxy that counts bytes, largest saving first.
1. Drop the browser, if your target lets you
This is the only change on the page worth eight times. Everything else is a rounding error next to it.
| Tool | Bytes down | Connections | Per 100,000 pages |
|---|---|---|---|
| curl_cffi | 59,744 | 1 | 6.0 GB |
| SeleniumBase | 425,130 | 5 | 42.5 GB |
| Patchright | 475,846 | 7 | 47.6 GB |
| Camoufox | 480,128 | 5 | 48.0 GB |
| Playwright | 482,162 | 7 | 48.2 GB |
8.07 times, same article. At a round dollar a gigabyte that is six dollars against forty-eight per hundred thousand pages. Use your own rate; the multiplier is the durable part.
from curl_cffi import requests
r = requests.get(url, proxies={"http": px, "https": px}, impersonate="chrome")
The extra 420 kilobytes is not waste from the site's point of view. It is images, fonts and an auth subdomain, and it is what makes the page render:
en.wikipedia.org 443,443
upload.wikimedia.org 31,039 images
auth.wikimedia.org 7,680
The trade is real. A tool that does not run JavaScript cannot scrape a page that builds itself in JavaScript. Check whether yours does before taking this saving, because it is the only one here that can break your scraper.
2. Block what you do not need
An ad blocker helps exactly in proportion to how much there is to block.
Camoufox ships uBlock Origin and enables it. On our Wikipedia article it came in at 480,128 bytes against plain Playwright's 482,162: under half a percent.
That is not a failure, it is the correct behaviour on a page with no trackers, and it is the reason to be careful with any bandwidth figure you read. We measured a real saving from the same blocker before, in how to use Camoufox with a proxy, where a page with trackers dropped from 35 requests to 27 because those requests never left.
So: a blocker pays on commercial pages and does nothing on clean ones, and it is available to any of these tools rather than being a reason to switch to one.
3. Turn off the browser's own chatter
Chrome talks to Google through your proxy before it loads your page. On a plain launch:
GET http://clients2.google.com/time/1/current
CONNECT accounts.google.com:443
CONNECT www.google.com:443
CONNECT example.com:443 <- the page you asked for
--disable-background-networking did not remove them in our test, so this is
worth verifying rather than assuming your flags worked. We also caught
SeleniumBase opening content-autofill.googleapis.com for 6,798 bytes during a
Wikipedia scrape.
Small per launch. Not small at ten thousand launches, and it is billed at the same per-gigabyte rate as the page you wanted.
4. Stop paying for failures
A blocked request is billed like any other. When we triggered a 403 by sending
requests with its default user agent, it cost 6,485 bytes for a 126-character
error page.
One header fixed that particular refusal:
| Request | Response |
|---|---|
requests, default user agent | 403, 126 characters |
requests, Chrome user agent | 200, 239,671 characters |
But do not over-learn it. The same swap did nothing against two sites refusing us through a bot-management product, which we go through in how to fix a 403 when scraping. Aggressive retries against something that is genuinely refusing you are a bandwidth line item with no page at the end of it.
Measuring your own
None of the numbers above will match your target. The method transfers and the figures do not.
Point your scraper at a proxy that counts bytes and read the total. Ours is in the repository with the test scripts, so you can run your own page through it before committing to a tool.
Failing that, a provider usage dashboard works: run a hundred pages, read the gigabytes, multiply.
Measuring your own, in code
The whole method is a proxy that adds up bytes per connection. This is the part that counts them:
def pipe_count(client, upstream, target):
"""Shuttle bytes both ways, totalling each direction."""
up_b = down_b = 0
while True:
ready, _, _ = select.select([client, upstream], [], [], 25)
if not ready:
break
for s in ready:
data = s.recv(65535)
if not data:
raise StopIteration
if s is client:
up_b += len(data)
upstream.sendall(data)
else:
down_b += len(data)
client.sendall(data)
log({"target": target, "up": up_b, "down": down_b})
Point a tool at it and read the total:
import json, collections
rows = [json.loads(l) for l in open("/tmp/bytes.jsonl")]
total = sum(r["down"] for r in rows)
by_host = collections.Counter()
for r in rows:
by_host[r["target"]] += r["down"]
print(f"{total:,} bytes over {len(rows)} connections")
for host, n in by_host.most_common():
print(f" {n:>8,} {host}")
Against the article measured above:
482,162 bytes over 7 connections
443,443 en.wikipedia.org
31,039 upload.wikimedia.org
7,680 auth.wikimedia.org
If you are sizing a budget rather than choosing a tool, the cost calculator takes a request count and an average page weight and gives you the monthly figure.
Limitations
One page, one day, one machine. A Wikipedia article is the friendly end of the spectrum: no advertising, no tracking, few third-party assets. A commercial page would raise every browser row, widen the gap to the HTTP clients, and narrow the gap between browsers by giving the blocker something to do. We measured neither.
We did not test whether any of these tools is blocked at volume, which is the other half of the cost question and needs many addresses to answer honestly.
We do not rank providers by anything except published price and measured behaviour, and no affiliate relationship enters a sort. Per-gigabyte pricing across the market is on the residential proxy page with the date each figure was read.
Sources
- curl_cffi PyPI, read 28 Aug 2026
curl_cffi 0.16.2
Questions
How much bandwidth does a headless browser use per page?
On the article we metered, between 425 and 482 kilobytes depending on the tool, against 60 kilobytes for an HTTP client fetching the same page. That is 48 gigabytes per hundred thousand pages against six.
What is the biggest saving available?
Dropping the browser, where you can. It was eight times on our test page and nothing else we measured came close. It only applies if your target renders server-side, which is the whole trade.
Does an ad blocker reduce proxy bandwidth?
In proportion to what there is to block. Camoufox ships uBlock Origin and saved under half a percent on a Wikipedia article, which carries no trackers. The saving we measured previously came from tracker requests never leaving, so a commercial page is where it pays and a clean one is where it does not.
Do blocked requests still cost bandwidth?
Yes. A 403 we triggered cost 6,485 bytes through the proxy for a 126-character error page. Failures are billed at the same rate as successes, which is worth remembering if you are retrying aggressively against something that is refusing you.
Which tool uses the fewest connections?
The HTTP clients, at one each. The browsers opened five to seven and contacted three or four hosts for a single article, including image and authentication subdomains nobody asks for by name.
