Tutorial
scrapy-playwright and Cloudflare Turnstile: JS-Rendered Forms
When the Cloudflare Turnstile widget only exists after JavaScript runs, let scrapy-playwright render the page, then solve the token, fill the form and parse.
By ZeroCaptcha Engineering5 min readPublished
Use scrapy-playwright for a Cloudflare Turnstile form when the widget only exists after
JavaScript runs: its sitekey is not in the HTML Scrapy downloads, or the page reads the token
through a JavaScript callback rather than the form field. Request the page with
meta={"playwright": True, "playwright_include_page": True}, wait for the widget in the Playwright
page, read its sitekey, get a token from a solving API, write it into cf-turnstile-response,
call the page’s callback if it has one, submit, and close the page. If the sitekey is already in the
server’s HTML, plain Scrapy does the same job without a browser.
This tutorial shows the spider, the settings it needs, and how to keep browser pages from piling up. Crawl only sites you are allowed to automate: see responsible captcha automation.
Set up
pip install scrapy scrapy-playwright httpxplaywright install chromiumexport ZEROCAPTCHA_API=… # the API's base URLexport ZEROCAPTCHA_KEY=… # your API keyscrapy-playwright needs two settings: its download handler for the schemes you crawl, and Twisted’s
asyncio reactor. The spider below sets both in custom_settings, so it runs on its own.
The spider
Save this as search_spider.py. It uses Scrapy’s async start() method, added in Scrapy 2.13; on
older versions, Scrapy’s docs say to “define also a synchronous start_requests() method that
returns an iterable”.
import asyncioimport osimport uuid
import httpximport scrapy
API = os.environ["ZEROCAPTCHA_API"]KEY = os.environ["ZEROCAPTCHA_KEY"]
FILL_TOKEN = """([token, callbackName]) => { for (const input of document.querySelectorAll('[name="cf-turnstile-response"]')) input.value = token; const callback = callbackName && window[callbackName]; if (typeof callback === "function") callback(token);}"""
async def solve_turnstile(page_url, sitekey, action=None, cdata=None): task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}} # The widget's data-action and data-cdata go in metadata, only when it sets them: many sites # check both when they verify the token. if action: task["metadata"]["action"] = action if cdata: task["metadata"]["cdata"] = cdata async with httpx.AsyncClient(base_url=API, timeout=15) as client: # One Idempotency-Key per task: a retried create with it returns the same task. created = (await client.post("/createTask", json={"clientKey": KEY, "task": task}, headers={"Idempotency-Key": str(uuid.uuid4())})).json() if created["errorId"]: raise RuntimeError(f"createTask: {created['errorCode']}") for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most await asyncio.sleep(2) reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]}) result = reply.json() if result["errorId"]: raise RuntimeError(f"getTaskResult: {result['errorCode']}") if result["status"] == "ready": return result["solution"]["token"] raise TimeoutError("no token within 180 seconds")
class SearchSpider(scrapy.Spider): name = "search" custom_settings = { "DOWNLOAD_HANDLERS": { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", }, "TWISTED_REACTOR": "twisted.internet.asyncioreactor.AsyncioSelectorReactor", "PLAYWRIGHT_BROWSER_TYPE": "chromium", "PLAYWRIGHT_MAX_PAGES_PER_CONTEXT": 4, }
async def start(self): yield scrapy.Request( "https://shop.example.com/search", meta={"playwright": True, "playwright_include_page": True}, errback=self.close_page, )
async def parse(self, response): page = response.meta["playwright_page"] try: widget = await page.wait_for_selector("form#search [data-sitekey]", timeout=15_000) token = await solve_turnstile( page.url, await widget.get_attribute("data-sitekey"), await widget.get_attribute("data-action"), await widget.get_attribute("data-cdata"), ) await page.fill("form#search input[name=q]", "running shoes") await page.evaluate(FILL_TOKEN, [token, await widget.get_attribute("data-callback")]) await page.click("form#search button[type=submit]") await page.wait_for_selector(".result", timeout=30_000) html = await page.content() finally: await page.close()
for result in scrapy.Selector(text=html).css(".result"): yield { "title": result.css("h2::text").get(), "url": response.urljoin(result.css("a::attr(href)").get()), }
async def close_page(self, failure): page = failure.request.meta.get("playwright_page") if page is not None: await page.close()Run it with scrapy runspider search_spider.py -O results.json.
How it works
- The page comes back open. With
playwright_include_page, the response carries the live Playwright page inresponse.meta["playwright_page"], so the callback can wait, click and run scripts in it. The README is firm that you must close it: “Always close pages when finished”, with an errback for requests that fail before your callback runs. - Wait for the widget, not a fixed time.
wait_for_selectorreturns as soon as the widget’s element exists, which on a JavaScript-rendered page may be well after the first response. - Solve late. The token is valid for 300 seconds and works once, so the spider asks for it right before it submits, after the page has loaded. See Cloudflare Turnstile token expiry.
- Fill the field and call the callback. In a real browser, the widget does both on success.
Pages that pass an inline callback to
turnstile.renderneed that function called instead; see submit a Cloudflare Turnstile token. - Parse the final HTML with Scrapy.
page.content()gives the page after the form’s results have rendered, and aSelectorparses it as usual.
Settings that keep it stable
| Setting | Why |
|---|---|
PLAYWRIGHT_MAX_PAGES_PER_CONTEXT |
Caps open tabs. Each tab waiting on a token is a whole page in memory; start at 4 and raise it as your machine allows. |
CONCURRENT_REQUESTS |
Keep it close to the page cap: queued browser requests hold nothing useful while they wait. |
PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT |
In milliseconds. Raise it for slow sites, so navigation doesn’t fail while the token is still good. |
PLAYWRIGHT_LAUNCH_OPTIONS |
Such as {"headless": False} while you debug the selectors. |
Only requests with "playwright": True in their meta go through the browser. Everything else, such
as result pages that are plain HTML, still uses Scrapy’s fast downloader in the same spider.
When not to use a browser
A browser costs far more per page than Scrapy’s downloader. If the sitekey is in the HTML and the
form posts cf-turnstile-response, use Scrapy with form2request
instead. If every page answers “Just a moment…” with HTTP 403, the site shows a Cloudflare
challenge page, not a widget, and a token does not help:
a challenge task returns the cf_clearance cookie with the user agent it is bound to, earned through your proxy; see the Cloudflare WAF and 5-second challenge solver.
The Cloudflare Turnstile solver for Playwright page has
more Playwright examples, and the Cloudflare Turnstile solver page
covers every language.
Sources
- scrapy-playwright, README (checked 1 October 2026).
- Scrapy: spiders, for the async
start()method, “Added in version 2.13” (checked 1 October 2026). - Playwright for Python: Page, for
wait_for_selector,evaluateandcontent(checked 1 October 2026). - Cloudflare challenges: supported browsers (checked 1 October 2026).
- Cloudflare Turnstile: client-side rendering (checked 1 October 2026).
- Cloudflare: Error 403 (checked 1 October 2026).
The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.