Tutorial
Crawlee for Python and Cloudflare Turnstile: A Crawler Tutorial
Handle Cloudflare Turnstile forms in Crawlee for Python's PlaywrightCrawler: read the widget, solve the token, fill and submit, with timeouts that fit.
By ZeroCaptcha Engineering4 min readPublished
To handle a Cloudflare Turnstile form in Crawlee for Python, let PlaywrightCrawler load the
page, wait for the widget, read its data-sitekey, get a token from a solving API with an async
HTTP client, write it into the cf-turnstile-response field, and submit. Two settings need
changing: request_handler_timeout, which is one minute by default, and the crawler’s concurrency,
because every page waiting on a token is an open browser tab. This tutorial shows the complete
crawler.
It is the Python counterpart of the Crawlee (JavaScript) tutorial. Crawl only sites you are allowed to automate: see responsible captcha automation.
Set up
python -m pip install 'crawlee[all]' httpxplaywright install chromiumexport ZEROCAPTCHA_API=… # the API's base URLexport ZEROCAPTCHA_KEY=… # your API keycrawlee[all] is the install Crawlee’s quick start gives for PlaywrightCrawler; playwright install fetches the browser.
The crawler
Save this as crawler.py:
import asyncioimport osimport uuidfrom datetime import timedelta
import httpxfrom crawlee import ConcurrencySettingsfrom crawlee.crawlers import PlaywrightCrawler, PlaywrightCrawlingContext
API = os.environ["ZEROCAPTCHA_API"]KEY = os.environ["ZEROCAPTCHA_KEY"]
FILL_TOKEN = """([token, callbackName]) => { for (const input of document.querySelectorAll('[name="cf-turnstile-response"]')) input.value = token; const callback = callbackName && window[callbackName]; if (typeof callback === "function") callback(token);}"""
async def solve_turnstile(client, page_url, sitekey, action=None, cdata=None): task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}} # The widget's data-action and data-cdata go in metadata, only when it sets them: many sites # check both when they verify the token. if action: task["metadata"]["action"] = action if cdata: task["metadata"]["cdata"] = cdata # One Idempotency-Key per task: a retried create with it returns the same task. created = (await client.post("/createTask", json={"clientKey": KEY, "task": task}, headers={"Idempotency-Key": str(uuid.uuid4())})).json() if created["errorId"]: raise RuntimeError(f"createTask: {created['errorCode']}") for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most await asyncio.sleep(2) reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]}) result = reply.json() if result["errorId"]: raise RuntimeError(f"getTaskResult: {result['errorCode']}") if result["status"] == "ready": return result["solution"]["token"] raise TimeoutError("no token within 180 seconds")
async def main() -> None: api = httpx.AsyncClient(base_url=API, timeout=15) crawler = PlaywrightCrawler( request_handler_timeout=timedelta(seconds=240), concurrency_settings=ConcurrencySettings(max_concurrency=10), )
@crawler.router.default_handler async def request_handler(context: PlaywrightCrawlingContext) -> None: page = context.page widget = page.locator("form#search [data-sitekey]").first await widget.wait_for(state="attached", timeout=15_000)
token = await solve_turnstile( api, page.url, await widget.get_attribute("data-sitekey"), await widget.get_attribute("data-action"), await widget.get_attribute("data-cdata"), ) await page.fill("form#search input[name=q]", "running shoes") await page.evaluate(FILL_TOKEN, [token, await widget.get_attribute("data-callback")]) await page.click("form#search button[type=submit]") await page.wait_for_selector(".result", timeout=30_000)
for link in await page.locator(".result a").all(): await context.push_data( {"title": (await link.text_content() or "").strip(), "url": await link.get_attribute("href")} )
try: await crawler.run(["https://shop.example.com/search"]) finally: await api.aclose()
if __name__ == "__main__": asyncio.run(main())Run it with python crawler.py. The results land in storage/datasets/default.
The settings that matter
| Setting | Default | What to set, and why |
|---|---|---|
request_handler_timeout |
timedelta(minutes=1) |
240 seconds. A page load, a solve and a slow form submission can pass a minute, and a handler cut off mid-way wastes a token that was already paid for. |
concurrency_settings |
Crawlee’s defaults | ConcurrencySettings(max_concurrency=10). Each running handler holds a browser tab while it waits on its token. |
max_request_retries |
3 |
A failed handler is retried, and each retry solves again. Keep it low for form pages. |
retry_on_blocked |
True |
“If True, the crawler attempts to bypass bot protections automatically.” It deals with blocked pages, not with a Turnstile widget inside a form. |
max_session_rotations |
10 |
Sessions rotate on proxy errors and blocks; the rotations don’t count towards the retries. |
One httpx.AsyncClient serves every handler, so connections to the API are reused. For larger
crawls, add an asyncio.Semaphore around solve_turnstile and honour Retry-After on HTTP 429:
the httpx tutorial has a client that does, with safe
retries through an Idempotency-Key.
Why write the field instead of clicking the widget
In a person’s browser, the widget fills cf-turnstile-response when it passes and calls the page’s
callback. A browser under Playwright control may not pass: Cloudflare says automated browsers are
not supported for solving production challenges. Writing the API’s token into the field, and
calling the callback named in data-callback, does what the widget would have done. The token
works once, within 300 seconds, so the handler solves right before it submits. See
Cloudflare Turnstile token expiry.
When a browser is more than you need
If the sitekey is in the server’s HTML and the form posts cf-turnstile-response, an HTTP crawler is
far cheaper than a browser. Crawlee’s HTTP-based crawlers can read the sitekey with a CSS selector,
and the httpx tutorial shows the form post itself. If
every page is a Cloudflare challenge (“Just a moment…”, HTTP 403, cf-mitigated: challenge), a
token does not help:
a challenge task returns the cf_clearance cookie with the user agent it is bound to; see the Cloudflare WAF and 5-second challenge solver.
More Python examples are on the Cloudflare Turnstile solver for Python
page.
Sources
- Crawlee for Python: quick start (checked 1 October 2026).
- Crawlee for Python: BasicCrawler and BasicCrawlerOptions, for the defaults above (checked 1 October 2026).
- Crawlee for Python: scaling crawlers,
for
ConcurrencySettings(checked 1 October 2026). - Cloudflare challenges: supported browsers (checked 1 October 2026).
- Cloudflare Turnstile: server-side validation, for the token’s 300-second lifetime (checked 1 October 2026).
- Cloudflare: Error 403 (checked 1 October 2026).
The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.