Cloudflare's default settings changed on Tuesday. As of September 15, the company now blocks "mixed-use" AI crawlers — bots that bundle search indexing together with model training and agent retrieval — from any customer page that carries advertising. The change applies automatically to new Cloudflare customers, to new sites added by existing customers, and to free-tier customers who never touched their configuration. Paid customers who already set their own rules keep them. The deadline was set on July 1, giving the AI industry ten weeks of notice, and the interesting part of Tuesday is not who complied but who the policy was actually built to squeeze.

It is not OpenAI, Anthropic or Perplexity. Those companies already run separately identified crawlers for separate purposes, which means a publisher can block training without losing answer-engine citations. The bots that fail Cloudflare's test are Googlebot, Microsoft's Bingbot and Apple's Applebot — three crawlers that do search indexing and AI data collection under a single user agent. Each of the three offers an opt-out that lives in robots.txt rather than in a distinct bot: Google-Extended, Applebot-Extended, and, for Bing, a `noarchive` robots meta attribute. Those are voluntary directives. Cloudflare's position is that a voluntary directive is not a control, and it has built a network-level gate instead.

Cloudflare has been unusually direct about the target. In its July announcement the company said the "world's largest search engine" — it did not need to write the name — has access to roughly "2x more information" than other AI firms, because it is difficult to stay discoverable in search without also feeding the AI stack. Google disputes the framing, pointing to Google-Extended, which lets site owners exclude their content from Gemini Apps and Vertex API training without affecting Search inclusion. That defense has a gap Cloudflare is exploiting: Googlebot crawls for Search including AI Overviews and AI Mode, so opting out of "training" does not opt a publisher out of having its content summarized in the results page that used to send it traffic.

"Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge," Cloudflare co-founder and CEO Matthew Prince said in the July announcement, citing the crossover point where bot traffic overtook human traffic online — a milestone that arrived roughly a year ahead of forecasts. "We hope that our proposed default changes encourage mixed-use crawlers to separate out search from agent use and training," Prince said. Cloudflare's own framing of the ad-supported carve-out is blunter: the setting "ensures that content that drives revenue cannot be crawled without explicit permission of those content owners."

The leverage is real because of scale. W3Techs' June 2026 survey put Cloudflare in front of 23.4% of all websites, and the company routinely describes itself as sitting on more than 20% of global internet request traffic. A default setting at that layer is closer to a standard than a product feature. Cloudflare's own measurements also explain why "mixed-use" needed its own category: in May 2026 the company attributed 51.8% of AI crawler requests to training and another 35.7% to mixed purposes, with search-only accounting for just 9.3%. More than half of AI crawl traffic, by Cloudflare's count, is spent re-fetching pages that have not changed — bandwidth publishers pay for and get nothing back from.

Alongside the block, Cloudflare renamed its 2025 pay-per-crawl marketplace to Pay Per Use, shifting the billing event from fetch to use. The launch partners are Ceramic.ai, an API search company, and You.com, an agent-facing search engine; publishers who opt in get paid when their content shows up in a Ceramic result or when a You.com agent pulls premium content. Two partners is not a marketplace. Cloudflare declined to give The Register any uptake numbers for the original Pay Per Crawl, which is the number that would tell you whether the tollbooth works.

Why it matters

This is the second escalation in fourteen months. On July 1, 2025, Cloudflare became the first major infrastructure provider to block AI crawlers by default for new domains, and publishers welcomed it. "Cloudflare's tools provide a strong framework for a more equitable exchange, offering a path for both industries to grow and thrive together," Danielle Coffey, president and CEO of the News/Media Alliance, said at the time. "By valuing and protecting the rights of publishers, we're ensuring that they can continue to create the high-quality content that fuels AI innovation."

What changed in 2026 is the pressure point. The 2025 move let publishers block scrapers they were already willing to lose. The 2026 move forces a decision they have spent two years avoiding: whether to treat Google as an AI company. Every attempt to price AI access to content has stalled on the same problem — publishers cannot afford to disappear from search, and the companies that control search have no reason to unbundle voluntarily. Cloudflare's answer is to make unbundling the price of admission to the ad-supported web, enforced at the CDN rather than negotiated contract by contract.

The objection from the AI side is that search indexing becomes collateral damage. It is fair and avoidable: the fix is a separate, separately verifiable crawler, which three of the largest labs already run. The harder question is enforcement. Bot identification at the edge is cat-and-mouse, and a default that mislabels a legitimate search crawler costs a publisher traffic immediately while the training data gets scraped some other way anyway.

What to watch

Whether Google, Microsoft or Apple splits its crawler, and when — the only outcome that would make this policy a success on Cloudflare's own terms. Whether paid-plan publishers, who were not switched over automatically, opt in. Whether Pay Per Use signs a partner anyone has heard of, and whether Cloudflare publishes participation numbers. And whether search referrals to Cloudflare-fronted ad-supported sites move in the next two quarters — that will show up in webmaster dashboards long before any press release.

“Now that the majority of traffic on the Internet is non-human, we must go further and act faster so that a sustainable ecosystem can emerge.”
— Matthew Prince, Co-founder and CEO, Cloudflare
23.4%
Share of all websites fronted by Cloudflare
35.7%
AI crawler requests classed as mixed-purpose
51.8%
AI crawler requests attributed to training
50%+
AI crawl traffic re-fetching unchanged pages