Cloudflare Turnstile
Scraping a Site With Cloudflare Turnstile: An Authorized Pipeline
Collect data you are allowed to collect from pages behind Cloudflare Turnstile: when a solver fits, how to add it to a pipeline, and when to ask first.
3 min readPublished Updated
Data pipelines meet Cloudflare Turnstile when a site puts it in front of a search form, a login or a listing page. A solving API can supply the token, but whether you should use one is a question about the data, not the widget. This guide starts with that question, then shows how a solver fits into a pipeline that is fast, economical and polite to the site.
First: are you allowed to collect this?
A solving API is appropriate when the automation itself is permitted. Common cases:
- Your own site, where you are building monitoring, a migration or an export.
- A client’s or partner’s site, with their written permission, such as a portal you are contracted to integrate with.
- Public data you may collect, under the site’s terms and the law that applies to you, such as prices you are permitted to monitor or public records.
It is not appropriate for accounts you do not own, data behind a login you are not entitled to, or anything a site’s terms forbid. ZeroCaptcha’s Acceptable Use Policy sets these rules for every task. Staff add a domain to the blocklist by hand after a report, or when its owner asks to opt out, and tasks for a blocked domain are refused. When in doubt, ask the site’s owner: many offer an API or a data export that is simpler than any scraper. Responsible captcha automation goes deeper.
Where Cloudflare Turnstile sits in a scraper
Turnstile protects actions, usually a form submission. A typical flow:
- Load the page with the form, in your HTTP client or browser, and keep its cookies.
- Read the widget’s sitekey, and its action and cData if set: see Find a Cloudflare Turnstile sitekey.
- Create a task with the page URL and sitekey, and wait for the token.
- Submit the form with the token in
cf-turnstile-response, as the page would. - Continue with the session the site gives you. Many sites do not ask again for a while.
Step 5 matters for cost: if the site sets a session after one successful form, you need one token per session, not one per page.
Solve just in time
A Turnstile token works once and expires 300 seconds after it is issued. In a pipeline with queues, that is easy to break: tokens solved at the start of a batch are stale by the time the last request goes out. Create each task when its request is about to be sent, and submit as soon as the token is ready. Cloudflare Turnstile token expiry shows the pattern.
Pace your requests
Being polite to the site is also what keeps a pipeline working:
- Limit concurrency per site, and respect its
robots.txtand published rate limits. - Back off when the site answers 429 or 503.
- Cache pages you have already fetched.
- Identify yourself in the user agent where the site’s terms ask for it.
On the solving side, cap the number of tasks in flight and honor Retry-After:
Rate limits and concurrency has a worker-pool example.
Browser or HTTP client?
- An HTTP client is faster and cheaper to run. It works when you can reproduce the form post: the fields, the cookies and the token.
- A headless browser handles pages that build the request in JavaScript. You still solve the token through the API, then hand it to the page. The Playwright, Puppeteer and Selenium pages show how.
Proxies
Solve through your own proxy with TurnstileTask only when the site needs it, such as a site that
serves some regions only, and then submit through the same proxy. See
Solve Cloudflare Turnstile with a proxy.
Watch the cost
You pay only for tasks that return a token; failed tasks cost nothing. The costs that add up in
scraping are tokens that expire in a queue and retries that create duplicate tasks. Send an
Idempotency-Key with every createTask, solve at the pace you submit, and put a daily spend cap
on the pipeline’s key. The pricing page has an estimator for a month of tasks.
Challenge pages are different
If the site shows a full-page “Just a moment…” check before any content, that is a Cloudflare
challenge page, not a Turnstile widget. It sets a cf_clearance cookie instead of producing a
form token. ZeroCaptcha’s challenge task returns that cookie with its user agent: see
Cloudflare challenge page vs Cloudflare Turnstile. For widgets, start
at the Cloudflare Turnstile solver page.