Skip to content

Tutorial

Scrapy and Cloudflare Turnstile: Submit Forms Without a Browser

A Scrapy spider that reads a Cloudflare Turnstile sitekey from the HTML, gets a token from an API without blocking the crawl, and posts the form.

By 4 min readPublished Updated

Scrapy cannot run the Cloudflare Turnstile widget, because it does not execute JavaScript, but it does not need to. The widget’s sitekey is in the page’s HTML, a solving API turns the sitekey into a token, and the form accepts that token in its cf-turnstile-response field. The one thing to get right is waiting for the token without stalling the crawl: use an async def callback, the asyncio reactor and an async HTTP client. This tutorial builds that spider.

Only automate sites you are allowed to: your own, a client’s, or one whose terms permit it. See responsible captcha automation.

What the page looks like

A form protected by Cloudflare Turnstile carries the widget as a div with the sitekey, and the widget adds a hidden input when it runs in a browser:

<form id="search" action="/search" method="post">
<input name="q" />
<div class="cf-turnstile" data-sitekey="0x4AAAAAAAB1cD2eF3gH4iJ5" data-action="search"></div>
<button type="submit">Search</button>
</form>

Scrapy sees the div but never the token. If the sitekey is not in the HTML at all, the page renders the widget from JavaScript with turnstile.render(): read the sitekey from the script, or use a browser crawler such as Crawlee.

The spider

It needs Python 3.10 or later. Install Scrapy, httpx and form2request (pip install scrapy httpx form2request), set ZEROCAPTCHA_API and ZEROCAPTCHA_KEY, and save this as search_spider.py:

import asyncio
import os
import uuid
import httpx
import scrapy
from form2request import form2request
API = os.environ["ZEROCAPTCHA_API"]
KEY = os.environ["ZEROCAPTCHA_KEY"]
async def solve_turnstile(page_url, sitekey, action=None, cdata=None):
task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}}
# The widget's data-action and data-cdata go in metadata, only when it sets them: many sites
# check both when they verify the token.
if action:
task["metadata"]["action"] = action
if cdata:
task["metadata"]["cdata"] = cdata
async with httpx.AsyncClient(base_url=API, timeout=15) as client:
# One Idempotency-Key per task: a retried create with it returns the same task.
created = (await client.post("/createTask", json={"clientKey": KEY, "task": task},
headers={"Idempotency-Key": str(uuid.uuid4())})).json()
if created["errorId"]:
raise RuntimeError(f"createTask: {created['errorCode']}")
for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most
await asyncio.sleep(2)
reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]})
result = reply.json()
if result["errorId"]:
raise RuntimeError(f"getTaskResult: {result['errorCode']}")
if result["status"] == "ready":
return result["solution"]["token"]
raise TimeoutError("no token within 180 seconds")
class SearchSpider(scrapy.Spider):
name = "search"
start_urls = ["https://shop.example.com/search"]
custom_settings = {
"TWISTED_REACTOR": "twisted.internet.asyncioreactor.AsyncioSelectorReactor",
}
async def parse(self, response):
widget = response.css("form#search [data-sitekey]")
if not widget:
self.logger.warning("no Cloudflare Turnstile widget on %s", response.url)
return
token = await solve_turnstile(
response.url,
widget.attrib["data-sitekey"],
widget.attrib.get("data-action"),
widget.attrib.get("data-cdata"),
)
form = form2request(
response.css("form#search"),
{"q": "running shoes", "cf-turnstile-response": token},
)
yield form.to_scrapy(callback=self.parse_results)
def parse_results(self, response):
for result in response.css(".result"):
yield {
"title": result.css("h2::text").get(),
"url": response.urljoin(result.css("a::attr(href)").get()),
}

Run it with scrapy runspider search_spider.py -O results.json.

How it works

  • The asyncio reactor. Scrapy can run coroutine callbacks on Twisted’s asyncio reactor, which lets await call any asyncio library, httpx included. Scrapy uses this reactor by default since 2.13; the TWISTED_REACTOR setting states it for older projects too.
  • Only the callback waits. While parse awaits the token, Scrapy keeps downloading and parsing other requests. A blocking requests.post in the same place would freeze the whole crawl for the length of every solve.
  • form2request reads the form’s action, method and fields, including the empty cf-turnstile-response input the widget would fill, and the data you pass overrides them with your query and the token. Scrapy deprecated FormRequest.from_response() in favor of this library in 2.16; its docs now use form2request for the same job.
  • Cookies carry over. Scrapy’s cookie middleware sends the session cookies the form page set, so the site sees the submission come from the same visitor that loaded the form.

Keep the token fresh

A Cloudflare Turnstile token is valid for 300 seconds and for one verification. In Scrapy, that means:

  • Solve inside the callback that submits the form, as above, never in start_requests.
  • Keep CONCURRENT_REQUESTS and your download delay such that the FormRequest goes out soon after the token arrives. A deep queue of pending requests can hold it past its lifetime.
  • If the site answers the form with an error, the token has been spent. Solve again before the retry. Cloudflare Turnstile token expiry explains the timeout-or-duplicate error the site sees otherwise.

Scaling it

Each coroutine waiting on a token holds nothing but a small object, so a spider can have dozens of solves in flight. To keep your read calls inside the API’s budget, open one httpx.AsyncClient for the whole spider (create it in open_spider of an extension or in the spider’s __init__, close it on spider_closed), cap solves in flight with an asyncio.Semaphore, and honor Retry-After. The httpx tutorial has that client, with safe retries through an Idempotency-Key.

When Scrapy is the wrong tool

Sources

The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.

Questions

Can Scrapy solve Cloudflare Turnstile on its own?

No. Scrapy does not run JavaScript, so the widget never produces a token. A spider can read the sitekey from the HTML, get a token from a solving API, and post it in the cf-turnstile-response field.

Does calling a CAPTCHA API from a Scrapy callback block the crawl?

Not if the callback is a coroutine and the API client is asynchronous. With Scrapy's asyncio reactor, await an httpx call and the other requests keep flowing while the token is prepared.

Does a Turnstile token help with a Cloudflare 'Just a moment' page in Scrapy?

No. That page is a Cloudflare challenge, not a Turnstile widget in a form. It is passed with a cf_clearance cookie, which is a different task.

Read next

This article is part of the Cloudflare Turnstile solver hub. Every task is charged only when a token is ready.

Get an API key