Tutorial
Crawlee and Cloudflare Turnstile: A PlaywrightCrawler Tutorial
Handle Cloudflare Turnstile forms in a Crawlee PlaywrightCrawler: read the sitekey, get a token from an API, fill the form, and set timeouts that fit.
By ZeroCaptcha Engineering4 min readPublished Updated
To handle a Cloudflare Turnstile form in Crawlee, let PlaywrightCrawler load the page, read
the widget’s data-sitekey, get a token from a solving API, write it into the
cf-turnstile-response input and submit. Two Crawlee settings need changing for it to work
reliably: the request handler’s 60-second default timeout, and how many pages solve at once. This
tutorial shows the whole crawler.
Crawl only sites you are allowed to automate. See responsible captcha automation.
Set up
npm install crawlee playwrightnpx playwright install chromiumexport ZEROCAPTCHA_API=… # the API's base URLexport ZEROCAPTCHA_KEY=… # your API keyNode.js 20 or later has fetch built in, which is all the API needs. Save the crawler below as
crawler.mjs.
The crawler
import { PlaywrightCrawler } from "crawlee";
const API = process.env.ZEROCAPTCHA_API;const KEY = process.env.ZEROCAPTCHA_KEY;const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));
async function call(method, body, headers = {}) { const response = await fetch(`${API}/${method}`, { method: "POST", headers: { "Content-Type": "application/json", ...headers }, body: JSON.stringify({ clientKey: KEY, ...body }), signal: AbortSignal.timeout(15_000), }); const reply = await response.json(); if (reply.errorId) throw new Error(`${method}: ${reply.errorCode}`); return reply;}
async function solveTurnstile(websiteURL, websiteKey, action, cdata) { // The widget's data-action and data-cdata go in metadata, only when it sets them: many sites // check both when they verify the token. const task = { type: "TurnstileTaskProxyless", websiteURL, websiteKey, metadata: {} }; if (action) task.metadata.action = action; if (cdata) task.metadata.cdata = cdata; // One Idempotency-Key per task: a retried create with it returns the same task. const { taskId } = await call("createTask", { task }, { "Idempotency-Key": crypto.randomUUID() }); const deadline = Date.now() + 180_000; while (Date.now() < deadline) { await sleep(2_000); const result = await call("getTaskResult", { taskId }); if (result.status === "ready") return result.solution.token; } throw new Error(`task ${taskId}: no token within 180 seconds`);}
const crawler = new PlaywrightCrawler({ maxConcurrency: 10, requestHandlerTimeoutSecs: 240, async requestHandler({ page, request, pushData, log }) { const widget = page.locator("form#search [data-sitekey]").first(); if ((await widget.count()) === 0) { log.warning(`No Cloudflare Turnstile widget on ${request.url}`); return; } const token = await solveTurnstile( page.url(), await widget.getAttribute("data-sitekey"), (await widget.getAttribute("data-action")) ?? undefined, (await widget.getAttribute("data-cdata")) ?? undefined, );
await page.fill("form#search input[name=q]", "running shoes"); await page.evaluate((value) => { for (const input of document.querySelectorAll('[name="cf-turnstile-response"]')) { input.value = value; } }, token); await page.click("form#search button[type=submit]"); await page.waitForSelector(".result", { timeout: 30_000 });
const results = await page.$$eval(".result a", (links) => links.map((link) => ({ title: link.textContent.trim(), url: link.href })), ); await pushData(results); },});
await crawler.run(["https://shop.example.com/search"]);Run it with node crawler.mjs. The results land in storage/datasets/default.
The settings that matter
requestHandlerTimeoutSecs: 240. Crawlee stops a request handler after 60 seconds by default. Solving a token usually takes seconds, but add page loads, a slow site and the form submission, and a busy minute is possible. A timeout mid-handler wastes a token that was already paid for, so give the handler room.maxConcurrency: 10. Each open page is a browser tab. Ten tabs, each waiting on a token, cost far more memory than ten HTTP requests. Raise it only as far as your machine allows.- Solve inside the handler. A Cloudflare Turnstile token is valid for 300 seconds and works once. Getting it in the same handler that submits the form keeps the gap to seconds. See Cloudflare Turnstile token expiry.
Why write the input instead of clicking the widget
In a real browser the widget fills cf-turnstile-response itself, when it passes. An automated
browser may not pass, or may be shown an interactive check. Writing the token from the API into
the input is the same thing the widget would have done. If the page reads the token through a
JavaScript callback instead of the input, it will not see your value:
submit a Cloudflare Turnstile token covers that case.
Crawlee’s blocking signals
Crawlee treats HTTP 401, 403 and 429 as blocked by default and retires the session that received
them. A Cloudflare challenge page (“Just a moment…”) is served with HTTP 403 and a
cf-mitigated: challenge header, so a crawler that meets one will rotate sessions and retry. That is a different problem from a
Cloudflare Turnstile form:
- A Turnstile widget in a form needs a token, as above.
- A challenge page on every request needs a
cf_clearancecookie. See the cf_clearance cookie explained and the Cloudflare WAF and 5-second challenge solver. - A 1020 or 1015 error page is a firewall rule or a rate limit, which neither fixes. See Cloudflare error 1020.
Crawlee also has a retryOnBlocked option, which tries to get past detected bot-protection pages.
It does not fill a Turnstile widget inside a form you submit.
A lighter crawler when the sitekey is in the HTML
If the form’s sitekey is in the server-rendered HTML, CheerioCrawler is much cheaper than a
browser. Read the sitekey with $("[data-sitekey]").attr("data-sitekey"), get the token with the
same solveTurnstile, and post the form with the handler’s sendRequest, which keeps the
session’s cookies. The got-scraping tutorial shows the
form post itself.
Sources
- Crawlee: PlaywrightCrawlerOptions,
for
requestHandlerTimeoutSecs(default 60) (checked 1 October 2026). - Crawlee: SessionPoolOptions,
for
blockedStatusCodes(default 401, 403 and 429) (checked 1 October 2026). - Crawlee: BasicCrawlerOptions,
for
retryOnBlocked(checked 1 October 2026). - Cloudflare challenges: detect a challenge response,
for the
cf-mitigatedheader, and Error 403, for the status of a challenge (checked 1 October 2026). - Cloudflare Turnstile: server-side validation (checked 1 October 2026).
The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.