Cloudflare Blocks Googlebot from Indexing and AI Training

Cloudflare's AI bot protection setting will automatically block Googlebot and Bingbot from September 15, 2026, removing sites from search results.

By Central
Cloudflare's new policy combines AI training and search indexing blocks, requiring proactive opt-out by September 15.
Highlights
  • Cloudflare's AI bot blocking will also block Googlebot and Bingbot starting September 15, 2026.
  • The block operates at the network edge, unlike robots.txt which is only a directive.
  • Site owners must opt out before the deadline to maintain search visibility.

Cloudflare blocks Googlebot from indexing and AI training through the same switch that was originally designed to keep AI crawlers away from publisher content. Starting September 15, 2026, every site with Cloudflare’s AI bot blocking enabled will also block Googlebot and Bingbot, and the decision to reverse course must be made before that date. This is not a robots.txt suggestion. It is an edge-level access refusal that can remove a page from search results entirely, and it lands in the same month that OpenAI’s search index, Google Analytics, and Google’s lawsuit over scraped results are all rewriting the rules of web visibility.

Cloudflare’s AI Bot Block Becomes a Googlebot Block on September 15

Cloudflare has confirmed that it will begin treating Googlebot and Bingbot as crawlers that both index for search and gather AI training data. That classification changes what the AI bot blocking setting actually does. Any site currently using a setting designed to block AI training, including the older “Block AI bots” switch, will automatically extend that block to Google and Microsoft’s search crawlers.

The policy takes effect on September 15, 2026. Site owners can opt out before that date, and doing so will allow Googlebot and Bingbot to crawl normally. After the date, the block is the default for anyone who has ever enabled an AI bot blocking option, even if they enabled it months ago and have not touched the controls since.

The practical effect is severe: a page that cannot be crawled by Googlebot will not be indexed. It will not appear in search results. It may still be accessible to human visitors, but it becomes invisible to the largest search distribution channel on the web. The fact that one site owner has already noticed Googlebot being blocked, while the cause remains unconfirmed, suggests the rollout may not be perfectly uniform. That uncertainty makes a proactive review of Cloudflare settings all the more important.

What does Cloudflare’s “Block AI bots” setting do now?

Cloudflare’s “Block AI bots” setting is a blanket control that refuses access to a list of crawlers associated with AI model training. Starting September 15, 2026, that list will include Googlebot and Bingbot because Cloudflare considers those bots to be collecting both search data and AI training data. The setting now functions as a combined search and training block rather than a separate AI-only filter.

How does this differ from robots.txt?

Robots.txt is a directive that tells crawlers what they are allowed to request, but it is a polite instruction that crawlers can choose to ignore. Cloudflare’s block operates at the network edge. A blocked crawler cannot reach the origin server at all, so the refusal is enforced at the infrastructure level. Googlebot will receive a connection failure or access denial rather than a robots.txt rule, which means the crawl simply does not happen. This is why the change has the potential to affect indexing so directly.

Which settings trigger the Googlebot block?

Any setting that blocks AI training will trigger the block, including the old “Block AI bots” switch. Cloudflare has also created newer controls such as “Block AI scrapers and crawlers” with separate options for search engines and social networks. If a site uses a configuration that blocks training but allows search, that configuration will be overridden on September 15 unless the site owner chooses to exclude Googlebot and Bingbot.

Why This Change Goes Beyond Bot Management

Cloudflare has effectively aligned itself with a particular interpretation of how search engines use web publisher content. The company has already built tools that let publishers decide whether their content can be used for AI training, AI search, or both. Treating Googlebot as a dual-purpose crawler is the logical extension of that philosophy. Google has stated that its crawler uses web content to improve AI models as well as to build its search index, so Cloudflare’s decision is grounded in the way Googlebot actually behaves.

For publishers, this is a strategic moment rather than a simple technical update. The setting that seemed harmless to enable in 2025, when the primary concern was blocking GPTBot or CCBot, now carries a much larger consequence. A publisher who wanted to keep content out of AI training will also be keeping it out of Google search unless they explicitly allow Googlebot through.

The decision creates a binary choice: allow Google access to everything it collects, including training data, or lose search visibility. There is no middle position in Cloudflare’s current model. That distinction will feel particularly sharp for websites that have built their traffic around SEO. For them, blocking Googlebot for any reason is a serious threat to revenue, even if the intent was to protect content from unauthorized AI use.

ChatGPT’s Search Index Was Already Serving Sites Without Content Deals

In the same week, new research from Resoneo, a French consultancy, showed that OpenAI’s in-house search index does not reserve good results for publishers with licensing agreements. The study examined ChatGPT answers captured in July and used a traffic field inside the answer data that revealed where each result came from. When a result was served from OpenAI’s own index, the answer was identical whether the publisher had a content deal with OpenAI or not.

The finding undermines a common assumption in publishing: that the only way to gain visibility in ChatGPT is to sign a content partnership. The data suggests the index itself is indifferent to commercial agreements. What matters is whether the page has been crawled, recorded, and retrieved. Deal status may affect other things, such as prominence or app integrations, but it does not appear to be the price of admission for the index behind most free-account answers.

OpenAI removed the provenance field from ChatGPT traffic in late July, which means the exact measurement cannot be repeated. The finding is now a snapshot rather than an ongoing experiment. Still, it changes the questions that publishers should be asking. Instead of “How do I get a deal with OpenAI?”, the more useful question is “What does OpenAI’s index actually store from my page?”

What does OpenAI’s index keep from a page?

According to Resoneo’s analysis, OpenAI’s index keeps a page’s title and roughly 200 characters from the beginning of the visible content. If a page template loads navigation, cookie banners, promotional bars, or advertisements before the main article, those elements consume the limited space that could otherwise represent the page’s core message.

For publishers, this has an immediate technical implication. The content that appears at the very top of the HTML, or at the very top of the rendered page, is the content most likely to be captured by OpenAI’s index. Putting a generic site-wide notice before the article is not harmless. It can crowd out the headline or the opening sentence that would have helped the page gain visibility in ChatGPT’s answers.

What did the SEO community make of the finding?

Radu Stoian, Technical Director at Enhance Media, framed the result as a reminder that AI search is not primarily a language model problem. He wrote: “The hardest part of AI search may not be the LLM. It’s the search index (and the harness, but that’s a different discussion).” His point is that retrieval quality, freshness, coverage, and the ability to match a query to the right page depend on infrastructure and data management, not on the model’s ability to generate fluent text.

For SEOs, this suggests that technical fundamentals still matter. Clean site architecture, fast rendering, useful metadata, and content that places meaningful text early on the page are all more important than deal-making. A publisher who can get into OpenAI’s index without a commercial agreement has a way to compete, but the competition depends on how the page appears to a crawler.

Google Analytics Will Benchmark Campaigns Against Similar Businesses

Google has announced that its Ask Advisor agent will compare campaign performance in Google Analytics against anonymized averages from similar businesses. The feature is part of a broader set of AI updates across Google Ads and Analytics. For marketers, the promise is simple: instead of wondering whether a campaign’s click-through rate or conversion rate is healthy, the platform will provide a reference point based on comparable accounts.

Google Analytics already has a peer group benchmarking feature. It places a property into a group built from the site’s industry category and other signals, and site owners can customize that group. Google has not said whether Ask Advisor will use the existing peer group, build its own comparison set, or offer a different level of control. The announcement also does not include a timeline for when the feature will reach accounts.

The distinction matters because benchmarks are only useful when the comparison group is truly similar. A small e-commerce site in home goods should not be measured against a national marketplace with a massive ad budget. The existing Analytics peer group gives property owners the ability to influence their grouping, which offers more control than the anonymized reporting available in Google Merchant Center. If Ask Advisor uses its own hidden group, marketers will lose that transparency.

What does this mean for campaign reporting?

Marketers who rely on Google Analytics to prove the value of their work will be able to add another layer of context to their reports. A campaign that is underperforming on a raw basis may actually be above average for the industry, while a campaign that looks strong may be lagging behind the peer set. The value of Ask Advisor will depend on how accurately Google can match accounts with similar businesses without exposing identifiable data.

What are the risks of AI-powered benchmarking?

Maryam Safari, Online Marketing Manager at PubliCare, tested the new feature by switching her GA4 account to English to trigger the beta, but she still saw the older Analytics Intelligence panel. Her conclusion was that language alone is not the gating factor in a gradual rollout. She also warned about what the agent is built on: “An agent sitting on top of broken tracking just produces the wrong answers faster.”

That warning is a useful reminder for anyone who is excited about AI reporting. Google Analytics is only as accurate as the tracking implementation beneath it. If events are mislabeled, conversions are double-counted, or sessions are attributed to the wrong campaigns, a benchmark tool will simply compare inaccurate data to anonymized averages. The result may look authoritative, but it is still garbage in, garbage out.

Google Refiles SerpApi Lawsuit With Licensing Terms Attached

Google has amended its DMCA complaint against SerpApi, adding terms from its content licensing agreements to the claims. The move is a direct response to a federal judge’s dismissal of Google’s original claims. Last month, the court dismissed both claims because Google had not demonstrated that copyright owners authorized its anti-scraping system. By including licensing terms, Google is trying to fill that gap and give its legal argument a contractual foundation.

Importantly, the claims related to search results without copyrighted content have been permanently dismissed. That limits the scope of the amended case. Google is no longer arguing that every search result snippet is protected by copyright. Instead, it is focusing on results that draw product images, reviews, and other licensed content found through its agreements with publishers and merchants.

Why is the SerpApi case important for SEO tools?

The case is part of a broader fight over who can collect Google search results at scale. Rank trackers, SERP monitors, and AI visibility tools all depend on the data that appears in search results. If Google wins, that data becomes harder to resell, and many third-party SEO products could be forced to change the way they collect information. If SerpApi prevails, the market for scraped search data remains open, at least for results that do not contain copyrighted content.

SerpApi’s response is the next filing, and nothing is final yet. The amendment is a significant step because it shows Google is listening to the court’s reasoning and adjusting its legal theory. It also shows that the question of who owns search results, and what parts of those results are protectable, is far from settled.

The Same Week Changed Access to Crawling, Indexing, and Search Data

These three stories are not isolated. They form a single pattern: access to the web is being renegotiated at three different layers. Cloudflare is deciding which crawlers can reach a site. OpenAI is deciding which parts of a page enter its index. A court is deciding who can collect Google’s search results at scale. Each decision affects how content moves through the modern web, but none of them is fully controlled by the publisher.

The old tools for managing access, robots.txt and contracts, feel inadequate now. Robots.txt only works if a crawler chooses to respect it. Contracts only cover the parties who sign them. Cloudflare’s edge control is more enforceable, but it cuts both ways: the same tool that protects a site from AI training can also remove it from search. OpenAI’s index shows that a deal is not the only route to visibility, but the hidden mechanics of that index leave publishers with little ability to know what is being stored. And the SerpApi lawsuit shows how quickly the legal ground can shift under the tools that measure visibility in the first place.

The only story in this list that arrives as a feature rather than a rule being rewritten is Google Analytics benchmarking. Campaign comparisons are a useful addition, but they depend on the same trust in tracking and anonymization that has become increasingly difficult to secure. The other three changes are structural. They change what can be crawled, what can be indexed, and who can resell the results.

The next few months will force publishers, SEO professionals, and analytics teams to make decisions they have never had to make before. Enabling an AI content protection tool now carries a search visibility cost. Building content for ChatGPT requires thinking about a 200-character window that has never appeared in SEO checklists. Monitoring search results may become legally riskier. None of these questions will be resolved by a single code update or court filing. But the direction is clear: the rules that governed crawling, indexing, and data collection for the past two decades no longer protect the people who depend on them.

Share This Article