Skip to content

Tutorial

Crawlee for Python and Cloudflare Turnstile: A Crawler Tutorial

Handle Cloudflare Turnstile forms in Crawlee for Python's PlaywrightCrawler: read the widget, solve the token, fill and submit, with timeouts that fit.

By 4 min readPublished

To handle a Cloudflare Turnstile form in Crawlee for Python, let PlaywrightCrawler load the page, wait for the widget, read its data-sitekey, get a token from a solving API with an async HTTP client, write it into the cf-turnstile-response field, and submit. Two settings need changing: request_handler_timeout, which is one minute by default, and the crawler’s concurrency, because every page waiting on a token is an open browser tab. This tutorial shows the complete crawler.

It is the Python counterpart of the Crawlee (JavaScript) tutorial. Crawl only sites you are allowed to automate: see responsible captcha automation.

Set up

Terminal window
python -m pip install 'crawlee[all]' httpx
playwright install chromium
export ZEROCAPTCHA_API=… # the API's base URL
export ZEROCAPTCHA_KEY=… # your API key

crawlee[all] is the install Crawlee’s quick start gives for PlaywrightCrawler; playwright install fetches the browser.

The crawler

Save this as crawler.py:

import asyncio
import os
import uuid
from datetime import timedelta
import httpx
from crawlee import ConcurrencySettings
from crawlee.crawlers import PlaywrightCrawler, PlaywrightCrawlingContext
API = os.environ["ZEROCAPTCHA_API"]
KEY = os.environ["ZEROCAPTCHA_KEY"]
FILL_TOKEN = """([token, callbackName]) => {
for (const input of document.querySelectorAll('[name="cf-turnstile-response"]')) input.value = token;
const callback = callbackName && window[callbackName];
if (typeof callback === "function") callback(token);
}"""
async def solve_turnstile(client, page_url, sitekey, action=None, cdata=None):
task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}}
# The widget's data-action and data-cdata go in metadata, only when it sets them: many sites
# check both when they verify the token.
if action:
task["metadata"]["action"] = action
if cdata:
task["metadata"]["cdata"] = cdata
# One Idempotency-Key per task: a retried create with it returns the same task.
created = (await client.post("/createTask", json={"clientKey": KEY, "task": task},
headers={"Idempotency-Key": str(uuid.uuid4())})).json()
if created["errorId"]:
raise RuntimeError(f"createTask: {created['errorCode']}")
for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most
await asyncio.sleep(2)
reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]})
result = reply.json()
if result["errorId"]:
raise RuntimeError(f"getTaskResult: {result['errorCode']}")
if result["status"] == "ready":
return result["solution"]["token"]
raise TimeoutError("no token within 180 seconds")
async def main() -> None:
api = httpx.AsyncClient(base_url=API, timeout=15)
crawler = PlaywrightCrawler(
request_handler_timeout=timedelta(seconds=240),
concurrency_settings=ConcurrencySettings(max_concurrency=10),
)
@crawler.router.default_handler
async def request_handler(context: PlaywrightCrawlingContext) -> None:
page = context.page
widget = page.locator("form#search [data-sitekey]").first
await widget.wait_for(state="attached", timeout=15_000)
token = await solve_turnstile(
api,
page.url,
await widget.get_attribute("data-sitekey"),
await widget.get_attribute("data-action"),
await widget.get_attribute("data-cdata"),
)
await page.fill("form#search input[name=q]", "running shoes")
await page.evaluate(FILL_TOKEN, [token, await widget.get_attribute("data-callback")])
await page.click("form#search button[type=submit]")
await page.wait_for_selector(".result", timeout=30_000)
for link in await page.locator(".result a").all():
await context.push_data(
{"title": (await link.text_content() or "").strip(), "url": await link.get_attribute("href")}
)
try:
await crawler.run(["https://shop.example.com/search"])
finally:
await api.aclose()
if __name__ == "__main__":
asyncio.run(main())

Run it with python crawler.py. The results land in storage/datasets/default.

The settings that matter

Setting Default What to set, and why
request_handler_timeout timedelta(minutes=1) 240 seconds. A page load, a solve and a slow form submission can pass a minute, and a handler cut off mid-way wastes a token that was already paid for.
concurrency_settings Crawlee’s defaults ConcurrencySettings(max_concurrency=10). Each running handler holds a browser tab while it waits on its token.
max_request_retries 3 A failed handler is retried, and each retry solves again. Keep it low for form pages.
retry_on_blocked True “If True, the crawler attempts to bypass bot protections automatically.” It deals with blocked pages, not with a Turnstile widget inside a form.
max_session_rotations 10 Sessions rotate on proxy errors and blocks; the rotations don’t count towards the retries.

One httpx.AsyncClient serves every handler, so connections to the API are reused. For larger crawls, add an asyncio.Semaphore around solve_turnstile and honour Retry-After on HTTP 429: the httpx tutorial has a client that does, with safe retries through an Idempotency-Key.

Why write the field instead of clicking the widget

In a person’s browser, the widget fills cf-turnstile-response when it passes and calls the page’s callback. A browser under Playwright control may not pass: Cloudflare says automated browsers are not supported for solving production challenges. Writing the API’s token into the field, and calling the callback named in data-callback, does what the widget would have done. The token works once, within 300 seconds, so the handler solves right before it submits. See Cloudflare Turnstile token expiry.

When a browser is more than you need

If the sitekey is in the server’s HTML and the form posts cf-turnstile-response, an HTTP crawler is far cheaper than a browser. Crawlee’s HTTP-based crawlers can read the sitekey with a CSS selector, and the httpx tutorial shows the form post itself. If every page is a Cloudflare challenge (“Just a moment…”, HTTP 403, cf-mitigated: challenge), a token does not help: a challenge task returns the cf_clearance cookie with the user agent it is bound to; see the Cloudflare WAF and 5-second challenge solver. More Python examples are on the Cloudflare Turnstile solver for Python page.

Sources

The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.

Questions

Why does my Crawlee for Python handler time out on a Cloudflare Turnstile page?

PlaywrightCrawler stops a request handler after one minute by default (request_handler_timeout). A page load, a solve and a form submission can take longer, so raise it, for example to timedelta(seconds=240).

Does Crawlee's retry_on_blocked solve Cloudflare Turnstile?

No. In Crawlee for Python it defaults to True and makes the crawler try to get past bot-protection pages. A Cloudflare Turnstile widget inside a form you submit still needs a token in its cf-turnstile-response field.

Is this the same as the JavaScript Crawlee tutorial?

The approach is the same, but the API differs: Python uses crawler.router.default_handler, a PlaywrightCrawlingContext, timedelta timeouts and ConcurrencySettings, and runs under asyncio.

Read next

This article is part of the Cloudflare Turnstile solver hub. Every task is charged only when a token is ready.

Get an API key