Capital & Compute

What a Crawl Is Worth: AI Bot Economics 2026

AI crawlers take up to 50,000 pages per visitor they send back. What that costs, what a toll would pay, and why blocking training now blocks Google.

· ai· search· economics· seo· monetization· By Capital & Compute

Across the first week of August 2025, Anthropic’s crawler fetched close to 50,000 pages from the average site for every one visitor it sent back. OpenAI’s fetched 887. Perplexity’s fetched 118. That ratio, not a ranking, is the number that decides what AI search is actually worth to a site.

Operator Crawls per referral, all verticals News and Publications Computer and Electronics
Anthropic ~50,000:1 2,500:1 8,800:1
OpenAI 887:1 152:1 401.7:1
Perplexity 118:1 32.7:1 88:1

All figures are Cloudflare Radar readings for the first week of August 2025, published in Cloudflare’s breakdown of AI crawler traffic by purpose and industry. The spread across the three columns is the part most coverage drops.

The ratio is the unit economics

A search engine and an AI crawler both cost a site the same thing: bandwidth and origin work. They differ entirely in what comes back. A search crawler fetches a page in order to send readers to it later. A training crawler fetches a page in order to absorb it, and the reader never arrives.

Cloudflare exposes this as the crawl-to-refer ratio: pages fetched divided by referrals returned. In Cloudflare’s 2025 Radar Year in Review, Google sat at 3:1 by mid-July after peaking near 30:1 in April, Microsoft cycled in a 50:1 to 70:1 band, Perplexity stayed below 200:1 from September onward, and Anthropic ranged from roughly 25,000:1 to 100,000:1 after May, touching 500,000:1 at its worst. Those are four different businesses, not four settings of one dial.

Crawl-to-refer ratio by operator and vertical, first week of August 2025A log-scale chart of three AI operators measured across three vertical selections in the same week of Cloudflare Radar data. Anthropic ranges from 2,500 crawls per referral on News and Publications sites to 8,800 on Computer and Electronics sites to nearly 50,000 across all verticals. OpenAI ranges from 152 to 401.7 to 887. Perplexity ranges from 32.7 to 88 to 118. A dashed break-even line sits at 10 crawls per referral, the point at which a toll of one tenth of a cent per crawl matches the value of a single referred visit worth one cent. All three operators sit far to the right of that line in every vertical.News and PublicationsAll verticalsComputer and Electronics10:1100:11,000:110,000:1100,000:1Pages crawled per referral returned (log scale)Break-even at $0.001/crawl vs a $0.01 visitAnthropicClaudeBot2.5k to 50kOpenAIGPTBot152 to 887PerplexityPerplexityBot33 to 118
Crawl-to-refer ratio by operator and vertical, first week of August 2025
OperatorNews and PublicationsComputer and ElectronicsAll verticals
Anthropic (ClaudeBot)2,500:18,800:150,000:1
OpenAI (GPTBot)152:1401.7:1887:1
Perplexity (PerplexityBot)32.7:188:1118:1
Each span is one operator measured three ways in the same week. The width of the span, not its position, is the finding: the vertical you are measured in changes the answer by up to 20x.Source: Cloudflare Radar, first week of August 2025

Your vertical changes the answer by 20x

The single most useful line in the Cloudflare data is the one nobody quotes: the ratio is not a property of the crawler. It is a property of the crawler and the content together.

Anthropic ran at 2,500:1 on News and Publications sites and close to 50,000:1 across all verticals in the same week. That is a twentyfold difference for one operator, measured at one moment. OpenAI moved from 152:1 to 887:1 across the same two selections, and Perplexity from 32.7:1 to 118:1. The direction is consistent: news gets referrals back, everything else largely does not.

The mechanism is not mysterious. News answers queries that recur, carry a freshness premium, and send a reader who wants the detail behind a summary. Reference and documentation content answers the query completely inside the assistant. A site whose pages fully satisfy the question they rank for is precisely the site that gets crawled hard and referred to rarely.

This is why a published industry-wide ratio is close to useless as a planning number, and why any post citing a single figure for “Anthropic” without naming the window and the vertical is quoting noise. Anthropic alone has been correctly cited at 2,500:1, 8,800:1, 50,000:1 and 500,000:1, all from the same source, all true, all of different things.

What a crawl costs to serve

The obvious answer is bandwidth, and the obvious answer is wrong, or at least incomplete. The real cost is that crawler traffic defeats the cache.

In Cloudflare’s April 2026 analysis of caching in the AI era, 32% of traffic across its network came from automated sources, at over 10 billion AI bot requests per week, roughly 80% of it crawler-focused. The request pattern is the problem. Human traffic concentrates: a small set of popular URLs absorbs most requests, which is exactly the shape a CDN is built for. AI crawlers do the opposite. Cloudflare describes them performing “sequential, complete scans” that reach “rarely visited or loosely related content,” with pages “once considered long-tail or rarely accessed” now frequently requested, and notes a measurable drop in cache hit rate once AI crawler traffic is added.

A cache miss is not a bandwidth event. It is an origin event: a request that reaches the application, runs whatever the application runs, and returns. For a static site that is cheap. For anything rendering per request, that is the expensive path, executed across the entire long tail of the URL space rather than the small hot set the capacity plan assumed.

The break-even: when a toll beats a referral

Here is the arithmetic that converts all of this into a decision, and it is simpler than the discourse around it suggests.

Cloudflare pay per crawl charges a flat, domain-wide price per successful retrieval, and the pricing documentation sets the minimum at $0.001 USD per crawl, billed on an HTTP 200. So N crawls are worth $0.001 × N as a toll. One referred visit is worth whatever your session actually earns, call it V.

Charging beats being referred when 0.001 × N exceeds V, which reduces to a break-even ratio of N = 1,000 × V.

Value of one referred visit Break-even crawl-to-refer ratio
$0.005 5:1
$0.01 10:1
$0.02 20:1
$0.05 50:1

Plug in your own session value rather than borrowing one. But note where the thresholds land. Even at five cents a visit, which is a strong display-advertising outcome, the break-even sits at 50:1. Every AI operator in the Cloudflare table clears that in every vertical, most of them by two or three orders of magnitude. Only Google, at 3:1 to 30:1 through 2025, is anywhere near the line.

That is the whole argument for tolls stated honestly: at the floor price, the referral traffic from AI crawlers is not a rounding error compared to what the same crawls would fetch as a toll. It is several orders of magnitude smaller.

The toll booth is mostly still closed

The obvious follow-up is to go and charge. This is where the coverage of this story is consistently misleading.

Pay per crawl was introduced in Cloudflare’s July 2025 announcement by Will Allen and Simon Newton, as a private beta. As of the AI Crawl Control documentation last updated 28 July 2026, it is still in closed beta, with a signup form and a note to contact your account executive if you are an enterprise customer. Headlines describing a live marketplace where any site can charge AI bots are describing a roadmap.

What does work today is the third option Cloudflare built into the same mechanism. A publisher can mark a crawler as charged even when that crawler has no billing relationship with Cloudflare. Cloudflare is explicit that this is “the functional equivalent of a network level block,” an HTTP 403 with no content returned, “but with the added benefit of telling the crawler there could be a relationship in the future.” A 402 with a price on it is a block that leaves a phone number.

So the practical 2026 decision is not what to charge. It is which of three categories to let through.

September 15, 2026: the default flips

On 1 July 2026, Cloudflare split AI bot management into three use cases: Search, Agent and Training. In its own definitions, Agent covers a bot “acting, usually in real time, on a person’s behalf,” including chat fetch bots and browser-use agents, with “often a human waiting on the other end.” Training covers a crawler whose fetch means “your data is permanently absorbed into the underlying architecture of the AI.”

In the same announcement, Cloudflare set a date. From 15 September 2026, for all new domains onboarding to Cloudflare, Training and Agent are blocked by default on pages that display ads, while Search remains allowed by default. The reasoning given is that “an ad is a signal that a website owner meant for a person to land there and see it,” so on those pages Cloudflare treats “human attention as the end goal.”

Two details matter more than the headline. The default applies to new domains onboarding to Cloudflare, not retroactively to every existing site, which several secondary write-ups have stated incorrectly. And it is ad-bearing pages specifically, which makes the trigger an inventory decision rather than a site-wide one.

The trap: blocking training now blocks Google

The second change landing on the same date is the one worth acting on, and it runs the other way.

Cloudflare states that multi-purpose crawlers, meaning those that combine Search with Training, “will be allowed/blocked according to all of their behaviors,” and that because “the defaults will be enforced by the most restrictive applicable rules, multi-purpose crawlers such as Googlebot, Applebot, and BingBot will be blocked by customers who have selected to block Training,” whether through the new controls or the legacy Block AI bots setting.

Read that again if you have ever ticked a box labelled “block AI training crawlers.” It now takes Googlebot with it.

Which crawlers separate their purposes, and which do notA three-lane classification of major AI crawlers by purpose. OpenAI, Anthropic and Perplexity each run separate user agents for search, assistant and training work, so a site owner can allow one and refuse another. Meta runs separate agents for assistant and training. Common Crawl is training only. Google, Microsoft and Apple run a single multi-purpose crawler that Cloudflare classifies as both search and training, so a rule blocking training also blocks the search crawler that sends traffic.Searchindexes, sends visitors backAgentfetching for a waiting humanTrainingabsorbed into the modelOpenAIOAI-SearchBotChatGPT-UserGPTBotAnthropicClaude-SearchBotClaude-UserClaudeBotPerplexityPerplexityBotPerplexity-UserMetaMeta-ExternalFetcherMeta-ExternalAgentCommon CrawlCCBotGoogleGooglebotGooglebotMicrosoftBingBotBingBotAppleApplebotApplebot
Which crawlers separate their purposes, and which do not
PrimitiveSearchAgentTraining
OpenAIOAI-SearchBotChatGPT-UserGPTBot
AnthropicClaude-SearchBotClaude-UserClaudeBot
PerplexityPerplexityBotPerplexity-Usernothing loads
Metanothing loadsMeta-ExternalFetcherMeta-ExternalAgent
Common Crawlnothing loadsnothing loadsCCBot
GoogleGooglebotnothing loadsGooglebot
MicrosoftBingBotnothing loadsBingBot
AppleApplebotnothing loadsApplebot
Rows with chips in two lanes are the problem. OpenAI, Anthropic and Perplexity ship separate user agents per purpose, so each is a separate switch. Google, Microsoft and Apple ship one crawler that does two jobs, so the most restrictive rule wins.Source: Cloudflare AI Crawl Control bot reference and the 1 July 2026 AI traffic options announcement

Cloudflare has been arguing this point at regulators rather than only at publishers. In its January 2026 submission on the UK Competition and Markets Authority consultation, it argued that requiring Google to split Googlebot by purpose “is not only technically feasible, but also a necessary and proportionate remedy,” after the CMA described separation as equally effective but declined to mandate it on burden grounds. Cloudflare’s own measurement in that submission put Googlebot at roughly 1.70x the unique URLs of ClaudeBot and 1.76x those of GPTBot.

By Cloudflare’s July 2026 bot report, 52% of crawler requests were for AI training as of June 2026, up from 22% in spring 2025, and more than a third of crawler activity still came from mixed-use bots that make crawl intent impossible to determine.

Why a flat price is the wrong mechanism

Assume the toll booth opens. The design that ships is a flat, domain-wide price per request, which means a one-sentence definition page and a six-month investigation are sold for the same tenth of a cent.

Pay-Per-Crawl Pricing for AI: The LM-Tree Agent, a preprint by Richard Archer, Soheil Ghili and Nima Haghpanah first posted to arXiv in April 2026 and revised in September, argues this is the central unsolved problem rather than a detail. Their framing is that content is too heterogeneous for a fixed pricing framework, and that the sub-types warranting different prices are too numerous to enumerate by hand.

Testing an adaptive pricing agent on 8,939 articles and 80,451 buyer queries from a major German technology publisher, with willingness to pay calibrated from actual AI crawler traffic, the preprint reports a 65% revenue gain over a single static price, a 47% gain over two-category pricing, and a 40% improvement on the publisher’s own eight-segment editorial taxonomy. This is a preprint in general economics, not a peer-reviewed result, and it evaluates one publisher’s library. Treat the direction as the finding rather than the magnitude.

The implication for anyone planning around tolls is that the floor price is not the price. It is the placeholder that exists because per-item pricing is hard, and the first mechanism that prices a page by what the page is worth will reprice the whole market.

What this site does

For disclosure, this site is a static publisher on Cloudflare Pages, so the arithmetic above is not neutral analysis of someone else’s problem.

The policy here is block training, allow answering, stated in the generated robots.txt as Content-Signal: search=yes, ai-input=yes, ai-train=no on every AI user agent, alongside an llms.txt index. Content Signals declares intended use rather than access, so it is a request rather than an enforcement mechanism, and it is worth being clear about the difference.

The reasoning is the ratio. Answer engines that cite are the only AI channel returning anything measurable, which is why the tooling for measuring it has become a paid software category with its own pricing problem. Training crawls return nothing by construction. That asymmetry, and what it does to the business model underneath, is the subject of why AI search broke the open web’s economics and the revenue strategies publishers are deploying in response.

Common questions

Frequently asked questions

What is a crawl-to-refer ratio?
The number of pages an operator fetches from a site divided by the number of visitors it sends back. Cloudflare publishes it per operator on Radar. A ratio of 887:1 means 887 pages were fetched for every one referral returned. Lower is better for the site.
How much can a site charge an AI crawler?
Cloudflare pay per crawl sets a minimum of $0.001 USD per successful crawl, charged on an HTTP 200 response, as a flat domain-wide price. As of the July 2026 documentation the feature is still in closed beta, so it is not available to most sites yet.
What changes on 15 September 2026?
Cloudflare sets new defaults for new domains onboarding to the platform: Training and Agent crawlers blocked by default on pages that display ads, Search allowed by default. Separately, multi-purpose crawlers are evaluated against all of their behaviors, so a rule blocking Training also blocks Googlebot, Applebot and BingBot.
Does blocking AI training crawlers hurt Google rankings?
It can, because Google does not separate its crawler by purpose. Cloudflare states that from 15 September 2026 the most restrictive applicable rule wins, so a site that has selected to block Training will block Googlebot, Applebot and BingBot too. Site owners can opt out in their Cloudflare Security settings.
Is it better to block AI crawlers or charge them?
At the $0.001 floor price, charging beats being referred above a break-even of 1,000 times the value of one visit, which is 10:1 for a visit worth a cent. Almost every AI operator sits far past that. In practice the charge option doubles as a block for crawlers with no billing relationship, so the two decisions are less separate than they appear.

Sources

Cloudflare (2025). A deeper look at AI crawlers: breaking down traffic by purpose and industry. Cloudflare blog, 28 August 2025. Primary, vendor-published network measurement. Source of all first-week-of-August-2025 crawl-to-refer figures. https://blog.cloudflare.com/ai-crawler-traffic-by-purpose-and-industry/ Verified 11 September 2026.

Cloudflare (2025). The 2025 Cloudflare Radar Year in Review. Cloudflare blog, December 2025. Primary, vendor-published. Source of the full-year 2025 ratio ranges for Google, Microsoft, Perplexity and Anthropic. https://blog.cloudflare.com/radar-2025-year-in-review/ Verified 11 September 2026.

Cloudflare (2026). Your site, your rules: new AI traffic options for all customers. Cloudflare blog, 1 July 2026. Primary, vendor announcement. Source of the Search, Agent and Training classification, the 15 September 2026 defaults, and the multi-purpose crawler rule. https://blog.cloudflare.com/content-independence-day-ai-options/ Verified 11 September 2026.

Cloudflare (2026). Content Independence Day, one year on: building the business model for the agentic Internet. Cloudflare blog, 1 July 2026. Primary, vendor-published. Source of the June 2026 crawl purpose shares and the mixed-use bot share. https://blog.cloudflare.com/agentic-internet-bot-report/ Verified 11 September 2026.

Cloudflare (2026). Why we are rethinking cache for the AI era. Cloudflare blog, 2 April 2026. Primary, vendor-published. Source of the automated traffic share, weekly AI bot request volume and the crawler request-pattern description. https://blog.cloudflare.com/rethinking-cache-ai-humans/ Verified 11 September 2026.

Cloudflare (2026). Cloudflare response to the UK CMA consultation on Google conduct requirements. Cloudflare blog, 30 January 2026. Primary, vendor position paper citing its own network measurement. https://blog.cloudflare.com/uk-google-ai-crawler-policy/ Verified 11 September 2026.

Allen, W. and Newton, S. (2025). Introducing pay per crawl: enabling content owners to charge AI crawlers for access. Cloudflare blog, 1 July 2025. Primary, vendor announcement. Source of the HTTP 402 mechanism, merchant-of-record arrangement and the charge-as-block behaviour. https://blog.cloudflare.com/introducing-pay-per-crawl/ Verified 11 September 2026.

Cloudflare (2026). Set a pay per crawl price and What is Pay Per Crawl?. Cloudflare AI Crawl Control documentation, last updated 28 July 2026. Vendor documentation. Source of the $0.001 minimum price and the closed-beta status. https://developers.cloudflare.com/ai-crawl-control/features/pay-per-crawl/ Verified 11 September 2026.

Cloudflare (2026). Bot reference. Cloudflare AI Crawl Control documentation. Vendor documentation. Source of the per-crawler operator and category classifications used in the three-lane figure. https://developers.cloudflare.com/ai-crawl-control/reference/bots/ Verified 11 September 2026.

Archer, R., Ghili, S. and Haghpanah, N. (2026). Pay-Per-Crawl Pricing for AI: The LM-Tree Agent. arXiv preprint 2604.01416 (econ.GN), submitted 1 April 2026, revised 11 September 2026. Preprint, not peer reviewed. https://arxiv.org/abs/2604.01416 Verified 11 September 2026.

Get each breakdown before it makes the rounds

You get one email when a new source-backed analysis goes live: what AI agents actually cost, which models are worth running, and what the benchmarks really mean. No hype.

No spam. Unsubscribe anytime.

← Back to AI markets