Tutorial
Scrapy and Cloudflare Turnstile: Submit Forms Without a Browser
A Scrapy spider that reads a Cloudflare Turnstile sitekey from the HTML, gets a token from an API without blocking the crawl, and posts the form.
By ZeroCaptcha Engineering4 min readPublished Updated
Scrapy cannot run the Cloudflare Turnstile widget, because it does not execute JavaScript, but it
does not need to. The widget’s sitekey is in the page’s HTML, a solving API turns the sitekey into
a token, and the form accepts that token in its cf-turnstile-response field. The one thing to
get right is waiting for the token without stalling the crawl: use an async def callback, the
asyncio reactor and an async HTTP client. This tutorial builds that spider.
Only automate sites you are allowed to: your own, a client’s, or one whose terms permit it. See responsible captcha automation.
What the page looks like
A form protected by Cloudflare Turnstile carries the widget as a div with the sitekey, and the
widget adds a hidden input when it runs in a browser:
<form id="search" action="/search" method="post"> <input name="q" /> <div class="cf-turnstile" data-sitekey="0x4AAAAAAAB1cD2eF3gH4iJ5" data-action="search"></div> <button type="submit">Search</button></form>Scrapy sees the div but never the token. If the sitekey is not in the HTML at all, the page
renders the widget from JavaScript with turnstile.render(): read the sitekey from the script, or
use a browser crawler such as Crawlee.
The spider
It needs Python 3.10 or later. Install Scrapy, httpx and form2request
(pip install scrapy httpx form2request), set ZEROCAPTCHA_API and ZEROCAPTCHA_KEY, and save
this as search_spider.py:
import asyncioimport osimport uuid
import httpximport scrapyfrom form2request import form2request
API = os.environ["ZEROCAPTCHA_API"]KEY = os.environ["ZEROCAPTCHA_KEY"]
async def solve_turnstile(page_url, sitekey, action=None, cdata=None): task = {"type": "TurnstileTaskProxyless", "websiteURL": page_url, "websiteKey": sitekey, "metadata": {}} # The widget's data-action and data-cdata go in metadata, only when it sets them: many sites # check both when they verify the token. if action: task["metadata"]["action"] = action if cdata: task["metadata"]["cdata"] = cdata async with httpx.AsyncClient(base_url=API, timeout=15) as client: # One Idempotency-Key per task: a retried create with it returns the same task. created = (await client.post("/createTask", json={"clientKey": KEY, "task": task}, headers={"Idempotency-Key": str(uuid.uuid4())})).json() if created["errorId"]: raise RuntimeError(f"createTask: {created['errorCode']}") for _ in range(90): # 90 polls, 2 seconds apart: 3 minutes at most await asyncio.sleep(2) reply = await client.post("/getTaskResult", json={"clientKey": KEY, "taskId": created["taskId"]}) result = reply.json() if result["errorId"]: raise RuntimeError(f"getTaskResult: {result['errorCode']}") if result["status"] == "ready": return result["solution"]["token"] raise TimeoutError("no token within 180 seconds")
class SearchSpider(scrapy.Spider): name = "search" start_urls = ["https://shop.example.com/search"] custom_settings = { "TWISTED_REACTOR": "twisted.internet.asyncioreactor.AsyncioSelectorReactor", }
async def parse(self, response): widget = response.css("form#search [data-sitekey]") if not widget: self.logger.warning("no Cloudflare Turnstile widget on %s", response.url) return token = await solve_turnstile( response.url, widget.attrib["data-sitekey"], widget.attrib.get("data-action"), widget.attrib.get("data-cdata"), ) form = form2request( response.css("form#search"), {"q": "running shoes", "cf-turnstile-response": token}, ) yield form.to_scrapy(callback=self.parse_results)
def parse_results(self, response): for result in response.css(".result"): yield { "title": result.css("h2::text").get(), "url": response.urljoin(result.css("a::attr(href)").get()), }Run it with scrapy runspider search_spider.py -O results.json.
How it works
- The asyncio reactor. Scrapy can run coroutine callbacks on Twisted’s asyncio reactor, which
lets
awaitcall any asyncio library, httpx included. Scrapy uses this reactor by default since 2.13; theTWISTED_REACTORsetting states it for older projects too. - Only the callback waits. While
parseawaits the token, Scrapy keeps downloading and parsing other requests. A blockingrequests.postin the same place would freeze the whole crawl for the length of every solve. form2requestreads the form’s action, method and fields, including the emptycf-turnstile-responseinput the widget would fill, and the data you pass overrides them with your query and the token. Scrapy deprecatedFormRequest.from_response()in favor of this library in 2.16; its docs now useform2requestfor the same job.- Cookies carry over. Scrapy’s cookie middleware sends the session cookies the form page set, so the site sees the submission come from the same visitor that loaded the form.
Keep the token fresh
A Cloudflare Turnstile token is valid for 300 seconds and for one verification. In Scrapy, that means:
- Solve inside the callback that submits the form, as above, never in
start_requests. - Keep
CONCURRENT_REQUESTSand your download delay such that theFormRequestgoes out soon after the token arrives. A deep queue of pending requests can hold it past its lifetime. - If the site answers the form with an error, the token has been spent. Solve again before the
retry. Cloudflare Turnstile token expiry explains the
timeout-or-duplicateerror the site sees otherwise.
Scaling it
Each coroutine waiting on a token holds nothing but a small object, so a spider can have dozens
of solves in flight. To keep your read calls inside the API’s budget, open one httpx.AsyncClient
for the whole spider (create it in open_spider of an extension or in the spider’s __init__,
close it on spider_closed), cap solves in flight with an asyncio.Semaphore, and honor
Retry-After. The httpx tutorial has that client, with
safe retries through an Idempotency-Key.
When Scrapy is the wrong tool
- The widget is rendered by JavaScript and the sitekey only exists at runtime: use a browser crawler, or read it from the bundle.
- Every page answers “Just a moment…” with HTTP 403 and a
cf-mitigated: challengeheader: that is a Cloudflare challenge page, not a Turnstile form. A token does not help. See Cloudflare challenge page vs Cloudflare Turnstile and the Cloudflare challenge solver. - The site answers error 1020 or 1015: those are firewall and rate-limit blocks. See Cloudflare error 1020 and error 1015.
Sources
- Scrapy: asyncio support and coroutines (checked 1 October 2026).
- Scrapy release notes, for 2.13 (the asyncio
reactor by default,
start()), 2.16 and 2.17 (FormRequest.from_response()deprecated in favor of form2request), and 2.19.0, the current release (checked 1 October 2026). - Scrapy: requests and responses, for form2request (checked 1 October 2026).
- Cloudflare Turnstile: client-side rendering, Cloudflare: detect a challenge response and Error 403 (checked 1 October 2026).
The team that builds and runs the ZeroCaptcha API. Articles are drafted with AI tools, then checked against the API's code and the primary sources each one cites.