Skip to content

Tutorial

scrapy-playwright and Cloudflare Turnstile: JS-Rendered Forms

When the Cloudflare Turnstile widget only exists after JavaScript runs, let scrapy-playwright render the page, then solve the token, fill the form and parse.

By 5 min readPublished

Use scrapy-playwright for a Cloudflare Turnstile form when the widget only exists after JavaScript runs: its sitekey is not in the HTML Scrapy downloads, or the page reads the token through a JavaScript callback rather than the form field. Request the page with meta={"playwright": True, "playwright_include_page": True}, wait for the widget in the Playwright page, read its sitekey, get a token from a solving API, write it into cf-turnstile-response, call the page’s callback if it has one, submit, and close the page. If the sitekey is already in the server’s HTML, plain Scrapy does the same job without a browser.

This tutorial shows the spider, the settings it needs, and how to keep browser pages from piling up. Crawl only sites you are allowed to automate: see responsible captcha automation.

Set up

Terminal window
pip install scrapy scrapy-playwright httpx
playwright install chromium
export ZEROCAPTCHA_API=… # the API's base URL
export ZEROCAPTCHA_KEY=… # your API key

scrapy-playwright needs two settings: its download handler for the schemes you crawl, and Twisted’s asyncio reactor. The spider below sets both in custom_settings, so it runs on its own.

The spider

Save this as search_spider.py. It uses Scrapy’s async start() method, added in Scrapy 2.13; on older versions, Scrapy’s docs say to “define also a synchronous start_requests() method that returns an iterable”.

import asyncio
import os
import uuid
import httpx
import scrapy
API = os.environ["ZEROCAPTCHA_API"]
KEY = os.environ["ZEROCAPTCHA_KEY"]
FILL_TOKEN = """([token, callbackName]) => {
for (const input of document.querySelectorAll('[name="cf-turnstile-response"]')) input.value = token;
const callback = callbackName && window[callbackName];
if (typeof callback === "function") callback(token);
}"""
async def solve_turnstile(page_url, sitekey, action=None, cdata=None):
task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}}
# The widget's data-action and data-cdata go in metadata, only when it sets them: many sites
# check both when they verify the token.
if action:
task["metadata"]["action"] = action
if cdata:
task["metadata"]["cdata"] = cdata
async with httpx.AsyncClient(base_url=API, timeout=15) as client:
# One Idempotency-Key per task: a retried create with it returns the same task.
created = (await client.post("/createTask", json={"clientKey": KEY, "task": task},
headers={"Idempotency-Key": str(uuid.uuid4())})).json()
if created["errorId"]:
raise RuntimeError(f"createTask: {created['errorCode']}")
for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most
await asyncio.sleep(2)
reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]})
result = reply.json()
if result["errorId"]:
raise RuntimeError(f"getTaskResult: {result['errorCode']}")
if result["status"] == "ready":
return result["solution"]["token"]
raise TimeoutError("no token within 180 seconds")
class SearchSpider(scrapy.Spider):
name = "search"
custom_settings = {
"DOWNLOAD_HANDLERS": {
"http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
"https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
},
"TWISTED_REACTOR": "twisted.internet.asyncioreactor.AsyncioSelectorReactor",
"PLAYWRIGHT_BROWSER_TYPE": "chromium",
"PLAYWRIGHT_MAX_PAGES_PER_CONTEXT": 4,
}
async def start(self):
yield scrapy.Request(
"https://shop.example.com/search",
meta={"playwright": True, "playwright_include_page": True},
errback=self.close_page,
)
async def parse(self, response):
page = response.meta["playwright_page"]
try:
widget = await page.wait_for_selector("form#search [data-sitekey]", timeout=15_000)
token = await solve_turnstile(
page.url,
await widget.get_attribute("data-sitekey"),
await widget.get_attribute("data-action"),
await widget.get_attribute("data-cdata"),
)
await page.fill("form#search input[name=q]", "running shoes")
await page.evaluate(FILL_TOKEN, [token, await widget.get_attribute("data-callback")])
await page.click("form#search button[type=submit]")
await page.wait_for_selector(".result", timeout=30_000)
html = await page.content()
finally:
await page.close()
for result in scrapy.Selector(text=html).css(".result"):
yield {
"title": result.css("h2::text").get(),
"url": response.urljoin(result.css("a::attr(href)").get()),
}
async def close_page(self, failure):
page = failure.request.meta.get("playwright_page")
if page is not None:
await page.close()

Run it with scrapy runspider search_spider.py -O results.json.

How it works

  • The page comes back open. With playwright_include_page, the response carries the live Playwright page in response.meta["playwright_page"], so the callback can wait, click and run scripts in it. The README is firm that you must close it: “Always close pages when finished”, with an errback for requests that fail before your callback runs.
  • Wait for the widget, not a fixed time. wait_for_selector returns as soon as the widget’s element exists, which on a JavaScript-rendered page may be well after the first response.
  • Solve late. The token is valid for 300 seconds and works once, so the spider asks for it right before it submits, after the page has loaded. See Cloudflare Turnstile token expiry.
  • Fill the field and call the callback. In a real browser, the widget does both on success. Pages that pass an inline callback to turnstile.render need that function called instead; see submit a Cloudflare Turnstile token.
  • Parse the final HTML with Scrapy. page.content() gives the page after the form’s results have rendered, and a Selector parses it as usual.

Settings that keep it stable

Setting Why
PLAYWRIGHT_MAX_PAGES_PER_CONTEXT Caps open tabs. Each tab waiting on a token is a whole page in memory; start at 4 and raise it as your machine allows.
CONCURRENT_REQUESTS Keep it close to the page cap: queued browser requests hold nothing useful while they wait.
PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT In milliseconds. Raise it for slow sites, so navigation doesn’t fail while the token is still good.
PLAYWRIGHT_LAUNCH_OPTIONS Such as {"headless": False} while you debug the selectors.

Only requests with "playwright": True in their meta go through the browser. Everything else, such as result pages that are plain HTML, still uses Scrapy’s fast downloader in the same spider.

When not to use a browser

A browser costs far more per page than Scrapy’s downloader. If the sitekey is in the HTML and the form posts cf-turnstile-response, use Scrapy with form2request instead. If every page answers “Just a moment…” with HTTP 403, the site shows a Cloudflare challenge page, not a widget, and a token does not help: a challenge task returns the cf_clearance cookie with the user agent it is bound to, earned through your proxy; see the Cloudflare WAF and 5-second challenge solver. The Cloudflare Turnstile solver for Playwright page has more Playwright examples, and the Cloudflare Turnstile solver page covers every language.

Sources

The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.

Questions

When do I need scrapy-playwright for a Cloudflare Turnstile form?

When the widget is rendered by JavaScript, so its sitekey is not in the HTML Scrapy downloads, or when the page reads the token through a JavaScript callback. If the sitekey is in the HTML and the form posts cf-turnstile-response, plain Scrapy is lighter.

Why do my scrapy-playwright pages leak memory?

A page requested with playwright_include_page stays open until you close it. Close it in a finally block in the callback, and in an errback for requests that fail, as the scrapy-playwright README recommends.

Does scrapy-playwright pass Cloudflare Turnstile by itself?

No. It renders pages in a Playwright browser, and Cloudflare says automated browsers are not supported for solving production challenges. Fill the widget's field with a token from a solving API instead.

Read next

This article is part of the Cloudflare Turnstile solver hub. Every task is charged only when a token is ready.

Get an API key