Troubleshooting
Cloudflare 403 Forbidden When Scraping: Which Cause Is It?
A Cloudflare 403 has five usual causes: a challenge page, a WAF block, a browser-signature ban, an IP or country ban, or the origin. How to tell them apart.
By ZeroCaptcha Engineering5 min readPublished
A 403 Forbidden from a site behind Cloudflare has five usual causes, and each needs a
different response: a challenge page (the response has the header cf-mitigated: challenge),
a WAF block (error 1020 in the body), a browser-signature ban (error 1010), an IP, network
or country ban (errors 1005 to 1009), or a 403 from the site’s own server (no Cloudflare
branding in the body). Only the first can be passed, by getting a cf_clearance cookie. The others
are decisions the site owner made, and the right response is to stop or to ask.
The rest of this article shows how to tell them apart from the response, with a script that does it for you. Automate only sites you are allowed to: see responsible captcha automation.
Where a Cloudflare 403 comes from
Cloudflare’s own 403 page (checked 1 October 2026) lists what makes it answer 403 “with Cloudflare branding”: WAF custom or managed rules “with challenge or block actions”, the Security Level setting, DDoS protection, Browser Integrity Check, validation checks, and “Most 1xxx Cloudflare error codes”. It also says that a 403 without Cloudflare branding comes from the origin server.
Two more facts make the diagnosis possible:
- Challenge pages announce themselves. Cloudflare sets
cf-mitigated: challengeon every challenge page, and “All challenge responses havetext/htmlas the content-type, even if you requested a different resource type”. - Blocks name their code in the body. “HTTP errors such as
409,530,403, and429are returned in the HTTP status header of a response, while 1XXX errors appear in the HTML body of the response.”
The five causes, side by side
| Cause | How it looks | Can it be passed? | Read more |
|---|---|---|---|
| Challenge page | 403, cf-mitigated: challenge, HTML titled “Just a moment…” |
Yes: pass it and send the cf_clearance cookie |
Cloudflare challenge types |
| WAF block | 403, “Error 1020: Access denied” in the body |
No | Cloudflare error 1020 |
| Browser signature | 403, error 1010 in the body |
No; the owner can turn off Browser Integrity Check | Cloudflare error 1010 |
| IP, network or country | 403, error 1005, 1006, 1007, 1008 or 1009 |
No | Errors 1006, 1007 and 1008, error 1009 |
| The origin’s own 403 | 403 without Cloudflare branding, often the site’s own error page |
Depends on the site: often a login or permission check | The site’s own documentation |
A 429 is a different status and a different problem, usually error 1015, a rate limit:
see Cloudflare error 1015.
Cloudflare’s error pages name these codes as: 1005 “Access Denied: Autonomous System Number (ASN) banned”, 1006, 1007 and 1008 “Access Denied: Your IP address has been banned”, 1009 “Access Denied: Country or region banned”, 1010 “The owner of this website has banned your access based on your browser’s signature”, and 1020 “Access denied”.
Tell them apart in code
This Python function reads a response and names the cause. It uses only what Cloudflare documents:
the cf-mitigated header, the Cf-Ray header Cloudflare adds to the responses it serves, and the
1xxx code in the body. The body pattern is a heuristic, since the error page’s exact markup is not
documented.
import re
import requests
ERROR_CODE = re.compile(r"(?:Error|error code:?)\s*(1\d{3})\b")CAUSES = { "1005": "network (ASN) ban", "1006": "IP ban", "1007": "IP ban", "1008": "IP ban", "1009": "country or region ban", "1010": "browser-signature ban (Browser Integrity Check)", "1020": "WAF block",}
def diagnose(response): if response.status_code != 403: return f"not a 403 ({response.status_code})" if response.headers.get("cf-mitigated") == "challenge": return "challenge page: pass it to get a cf_clearance cookie" match = ERROR_CODE.search(response.text) if match: code = match.group(1) return f"Cloudflare error {code}: {CAUSES.get(code, 'see Cloudflare 1xxx errors')}" if "cf-ray" in response.headers: return "403 served through Cloudflare with no 1xxx code: likely a WAF managed rule or the origin" return "403 from the origin server"
response = requests.get("https://shop.example.com/", timeout=30)print(diagnose(response))Run it once from the machine and proxy your crawler uses: the answer can differ by IP address, and by client. The same URL can challenge a script and serve a browser.
What to do about each
- A challenge page. Your client needs a
cf_clearancecookie. Cloudflare ties it to “the specific visitor and device it was issued to”; in practice, send it from the same IP address with the same user agent that earned it. A browser that runs the challenge can earn one; a plain HTTP client cannot. ZeroCaptcha’s challenge task passes the page through your own proxy and returns the cookie with the user agent it was issued for. See the cf_clearance cookie explained and the Cloudflare WAF and 5-second challenge solver. - A WAF block (1020). A rule matched something about the request: its path, country, headers, ASN or bot score. Nothing you send passes a block. Ask the site owner, who can find the event by its Ray ID, or stop.
- A browser-signature ban (1010). Browser Integrity Check judged the client by its headers. Cloudflare’s advice is to contact the site owner, who can turn it off. Sending a real browser’s full headers is how legitimate clients avoid it.
- An IP, network or country ban. The address is the problem: the owner banned it, its network or its country. See choosing proxies for how addresses are judged, and respect the ban on sites that don’t want your traffic.
- The origin’s 403. Cloudflare passed the request, and the site refused it. Look for a missing session, a login, a CSRF token or an API key, as you would without Cloudflare.
Why the same request gets 403 in a script and 200 in a browser
The usual answer is the first cause: the site challenges clients that don’t look like a browser,
and a browser passes the challenge silently while a script gets the 403. Cloudflare’s list of
supported browsers names command-line tools such as curl and wget, which run no JavaScript,
among the clients its challenges do not support. Clients also differ below HTTP: TLS fingerprints such as JA4
tell a Python library from Chrome before any header is read.
Sources
- Cloudflare: Error 403 (checked 1 October 2026).
- Cloudflare: 1xxx errors, including error 1020 and error 1010 (checked 1 October 2026).
- Cloudflare challenges: detect a challenge response (checked 1 October 2026).
- Cloudflare challenges: clearance (checked 1 October 2026).
- Cloudflare challenges: supported browsers (checked 1 October 2026).
- Cloudflare: HTTP headers,
for
Cf-Ray(checked 1 October 2026).
The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.