Cloudflare Enables Selective AI Bot Blocking for Search and Training

New granular controls from Cloudflare let website operators selectively block AI training bots while preserving search indexing access.

By Central
Cloudflare's new bot classification system distinguishes search, agent, and training use cases for independent management.
Highlights
  • Cloudflare's new system distinguishes three AI bot categories: Search, Agent, and Training, each managed independently.
  • Starting September 15, Cloudflare will block Agent and Training bots by default on pages displaying advertising.
  • Multi-purpose crawlers face the strictest applicable rule, losing access to ad-supported pages unless manually overridden.

Cloudflare has introduced more granular controls that allow website operators to selectively manage how artificial intelligence bots interact with their content. The new system makes it possible to permit bots that support search indexing while simultaneously blocking those that scrape data for training AI models.

This development addresses a growing tension that has come to define the modern web publishing landscape. Website operators face a fundamental dilemma: making content publicly available ensures visibility in search results and drives traffic, but it also exposes that content to AI providers who harvest it for model training without compensation or attribution. As AI-generated answers increasingly replace traditional search results, traffic that once flowed to publisher websites is instead captured by AI platforms.

Cloudflare now provides a mechanism to rebalance this dynamic. The company has introduced a categorization system that distinguishes between three distinct bot use cases, each of which can be managed independently.

Three Categories of AI Bot Behavior

The classification system reflects the different ways AI systems interact with web content. The first category, Search, covers bots that build search indexes. These bots create opportunities for websites to receive traffic from search engines, making them generally beneficial for publishers. The second category, Agent, describes bots that perform real-time actions on behalf of a user, such as filling out a form or completing a purchase. The third category, Training, is the use case that has generated the most concern among publishers. Bots in this category extract information to train new AI models or improve existing systems, often without any reciprocal value flowing back to the content creator.

What makes Cloudflare approach notably flexible is the ability to apply different rules to pages based on whether they display advertising. The underlying logic holds that pages with ads are intended to generate revenue through traffic, making them prime candidates for blocking training-oriented AI bots. Pages without ads may serve different purposes and can be treated differently under the policy.

New Default Standards Take Effect September 15

Beginning September 15, Cloudflare will enforce new default settings for all domains that join the platform. Under these defaults, the Agent and Training categories will be blocked on pages that display advertising. The Search category will remain permitted across all pages. This represents a significant shift in the baseline assumption about how AI bots should be treated.

For multi-purpose crawlers capable of fulfilling several use cases simultaneously, the strictest applicable rule will apply across all categories after September 15. A crawler that functions both as an Agent and for Search purposes will lose access to pages with advertising unless the website operator manually overrides the default configuration.

Content Usage Signals Add Another Layer of Control

Beyond the use-case classification, Cloudflare has introduced a second dimension that controls what crawlers can actually do with the content they retrieve. Three levels of permitted usage are defined. The first, immediate, covers pure interaction where nothing is retained or stored. The second, reference, allows indexing, excerpt generation, and link creation. The third, full, permits summarization and reproduction of content.

Cloudflare is also testing a new signal called use that can be specified in the robots.txt file. This signal allows website operators to declare the permitted usage level directly in their robots.txt directives. A typical entry might specify use=reference, indicating that a crawler may index and create excerpts but may not reproduce content in full.

Verified Bots Lose Automatic Access

Another important change affects verified bots. Previously, bots that could demonstrate their identity through cryptographic signatures, published IP ranges, or reverse DNS lookups were automatically granted access by default, provided they exhibited no abusive behavior and respected robots.txt directives. Cloudflare has now revoked this default permission. Verified status no longer guarantees access, giving website operators full control over which identified bots can interact with their content.

To support better identification and management of AI crawlers, Cloudflare has launched BotBase, a new database that catalogs all known bots. This resource provides website operators with comprehensive information about the bots attempting to access their sites, enabling more informed policy decisions.

What This Means for Publishers and Website Operators

The changes represent a practical response to what has become known as the content dilemma in the age of AI. Publishers have watched their content feed the training pipelines of large language models while receiving nothing in return, even as AI-generated summaries reduce the incentive for users to visit original sources. Cloudflare granular controls do not solve this problem entirely, but they give content creators a meaningful tool for asserting their preferences.

The system allows operators to continue benefiting from search engine crawlers that drive traffic while blocking the training-oriented crawlers that extract value without reciprocity. The ability to apply different rules based on advertising status adds a further layer of nuance, enabling operators to protect their monetized content while keeping non-commercial pages accessible.

Practical Implications for Multi-Purpose Crawlers

The treatment of multi-purpose crawlers deserves particular attention. As AI systems become more sophisticated, the distinction between a search crawler and a training crawler is not always clear-cut. Many major AI platforms operate crawlers that serve multiple functions simultaneously. Cloudflare approach of applying the strictest relevant rule ensures that operators do not inadvertently grant broad access to a crawler that claims a benign purpose while also collecting data for training.

This design choice reflects a conservative but sensible default. Website operators who wish to grant broader access can always do so through manual configuration, but the platform will not assume consent where ambiguity exists.

Context and Industry Significance

The timing of these changes is notable. The relationship between content publishers and AI platforms has become increasingly adversarial as the scale of AI training data collection has become apparent. Several high-profile lawsuits have been filed by publishers and content creators against AI companies, alleging unauthorized use of copyrighted material for training purposes. Cloudflare approach offers a technical rather than legal mechanism for managing this relationship.

By positioning itself as a gatekeeper that can enforce nuanced policies at the infrastructure level, Cloudflare is providing a service that its customers have increasingly demanded. The company is also shaping norms around acceptable AI crawler behavior, particularly through the introduction of the use signal in robots.txt, which could become an industry standard if widely adopted.

Strategic Considerations for Website Operators

For operators evaluating how to configure these controls, several factors merit consideration. The decision to block training bots on advertising-supported pages is straightforward for most commercial publishers, as the value extraction by AI platforms directly undermines the advertising revenue model. However, operators with different business models may reach different conclusions.

Non-commercial sites, educational resources, and open-access publications may choose to permit training access as a form of contribution to the broader AI ecosystem. The granularity of Cloudflare controls makes such selective permission feasible, allowing operators to distinguish between pages and purposes rather than applying a blanket policy.

The changes to verified bot access also warrant attention. Even well-known and reputable bots must now be explicitly permitted rather than automatically trusted. This represents a shift in the default security posture that aligns with the broader industry trend toward zero-trust models, where no entity is granted access without explicit authorization.

Looking Ahead

Cloudflare implementation of these controls, combined with the BotBase identification database and the use signal in robots.txt, provides a framework that could significantly influence how AI bots interact with the web. The September 15 deadline for new defaults gives current operators time to review their configurations and determine whether the new standards align with their preferences.

As AI companies continue to develop increasingly sophisticated crawlers and as the legal landscape around training data remains unsettled, technical controls of this kind offer a practical middle ground. They cannot replace legislation or licensing agreements, but they give content creators a direct mechanism for expressing their intent at the infrastructure level. For an industry searching for balance between openness and protection, that represents meaningful progress.

Share This Article