{"id":75084,"date":"2026-08-08T10:30:00","date_gmt":"2026-08-08T14:30:00","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=75084"},"modified":"2026-08-08T09:25:28","modified_gmt":"2026-08-08T13:25:28","slug":"cloudflare-ai-bot-blocking-googlebot","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/cloudflare-ai-bot-blocking-googlebot\/","title":{"rendered":"Cloudflare AI bot setting blocks Googlebot and Bingbot"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">For website owners and SEO professionals who rely on Cloudflare\u2019s security and performance suite, a troubling discovery has surfaced: the platform\u2019s new AI bot-blocking feature appears to be inadvertently blocking the industry\u2019s most important web crawlers. Several site administrators have reported that enabling Cloudflare\u2019s &#8220;AI Training&#8221; block, a feature introduced to prevent artificial intelligence companies from scraping content, is resulting in HTTP 403 errors for both Googlebot and Bingbot. This unintended consequence has immediate implications for <a href=\"https:\/\/overcentral.com\/en\/ai-search-visibility-citations\/\" title=\"AI Search Visibility: Citations Are Not Recommendations\" data-iacss-internal=\"1\">search visibility<\/a>, indexing, and traffic, raising urgent questions about how content delivery networks are implementing AI-era controls \u2014 and at what cost to the open web\u2019s foundational discovery mechanisms.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Cloudflare\u2019s AI bot-blocking feature inadvertently disrupts Googlebot and Bingbot access<\/h2>\n\n\n<p class=\"wp-block-paragraph\">In a recent discussion on Reddit, a website operator detailed a concerning scenario: after enabling Cloudflare&#8217;s new AI bot-blocking setting designed to restrict crawlers used for training large language models, they observed that both Googlebot and Bingbot began receiving 403 Forbidden responses when attempting to retrieve their sitemap. The user explicitly noted that this was not a test involving an imitated or spoofed Googlebot; the Cloudflare dashboard itself confirmed that these legitimate search engine crawlers were being blocked. When the AI Training block was subsequently disabled, the sitemap became accessible again without issue.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The user&#8217;s exact observation was: &#8220;When I set AI Training = Block, both Googlebot and Bingbot start receiving HTTP 403 responses when trying to fetch my sitemap. As soon as I disable the AI Training block, the sitemap is accessible again.&#8221;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What makes this incident particularly significant is that the <a href=\"https:\/\/overcentral.com\/en\/google-search-console-indexing-report-resumes\/\" title=\"Google Search Console Gets Indexing Report Data from July 24\" data-iacss-internal=\"1\">Google Search Console<\/a> for the affected domain began displaying error messages during the same timeframe, confirming that the indexing signals were not just theoretical \u2014 they had tangible, measurable consequences for search engine visibility.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Understanding the mechanism: how Cloudflare classifies and restricts bots<\/h2>\n\n\n<p class=\"wp-block-paragraph\">Cloudflare&#8217;s AI bot-blocking feature represents a broader trend among web infrastructure providers who are responding to the explosive growth of AI crawlers. With the proliferation of large language models such as OpenAI&#8217;s GPT, Google&#8217;s Gemini, and Anthropic&#8217;s Claude, the demand for training data has driven an unprecedented surge in automated scraping. Website operators have grown increasingly concerned about their intellectual property being harvested without compensation, attribution, or control, and Cloudflare&#8217;s response was to offer granular controls over which bots could access their clients&#8217; properties.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The AI Training setting applies the strictest possible restrictions to any crawler that is identified as being used for multiple purposes. Under Cloudflare&#8217;s policies, if a bot is employed for both search engine indexing and AI model training, it is subject to the most restrictive treatment automatically, unless the site owner explicitly overrides the default configuration. This approach, announced in early July, was scheduled to become the standard default starting September 15, affecting thousands of websites globally.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, the unexpected overlap with major search crawlers appears to stem from how Cloudflare classifies these bots. Since Googlebot and Bingbot are increasingly part of ecosystems that also involve AI training \u2014 Google&#8217;s crawlers feed its search index as well as its broader AI initiatives, while Microsoft&#8217;s Bing powers various AI products including Copilot \u2014 the automated classification system appears to be conflating the search indexing and AI training use cases.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What is the AI Training block and how does it work?<\/h2>\n\n\n<p class=\"wp-block-paragraph\">The AI Training block is a Cloudflare security setting that prevents automated crawlers and scrapers from accessing a website&#8217;s content when those crawlers are identified as being used to train artificial intelligence models. The setting applies the strictest access controls to any bot that engages in activities beyond traditional search indexing. When enabled, the block returns an HTTP 403 Forbidden response to any detected AI crawler, effectively denying access to pages, images, scripts, and sitemaps that might otherwise be used as training data.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The response gap: Cloudflare&#8217;s silence on the issue<\/h2>\n\n\n<p class=\"wp-block-paragraph\">As of the time of reporting, Cloudflare has not issued any official response to the reports of Googlebot and Bingbot being blocked by their AI Training setting. Neither the Reddit thread nor any public press release has addressed the incident directly. The company, known for its generally transparent and responsive approach to community concerns, has remained conspicuously quiet on this issue, leaving website owners to speculate about whether this is an isolated bug, an intentional design choice, or an unanticipated side effect of broader policy changes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">This silence is notable given Cloudflare&#8217;s stature as one of the world&#8217;s largest content delivery network providers, serving millions of internet properties. The absence of an official statement may reflect a need for internal investigation, but it also compounds the uncertainty for clients who rely on Cloudflare to protect their sites without inadvertently harming their SEO performance. In the absence of guidance, many site operators are faced with a difficult choice: enable AI protections and risk search engine blockers, or forgo those protections and leave their content vulnerable to unmanaged AI scraping.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Broader context: Cloudflare&#8217;s July announcement and the September deadline<\/h2>\n\n\n<p class=\"wp-block-paragraph\">In early July, Cloudflare announced a significant shift in its bot management policies. The company stated that starting September 15, it would apply the strictest available setting to any crawler that served dual purposes. This was presented as a proactive measure to give website owners more control over their content and to address growing unease about the volume of AI training data being harvested without consent.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Specifically, Cloudflare&#8217;s policy indicated that if a bot was being used for search indexing in addition to AI model training, it would be automatically restricted from accessing pages that contained advertising or that the site owner had designated as protected. Unless the owner manually selected an alternative configuration, the default was to treat these multi-purpose bots as subject to the most restrictive level of access. The announcement was widely viewed as a major development in the ongoing battle over content ownership, data rights, and the responsibilities of platform providers to protect their clients&#8217; intellectual property.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, the July announcement did not explicitly detail how Googlebot and Bingbot would be classified. Given that both <a href=\"https:\/\/www.google.com\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Google<\/a> and <a href=\"https:\/\/www.microsoft.com\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Microsoft<\/a> operate comprehensive AI research divisions and have incorporated AI training into their core product roadmaps, the classification of these bots as multi-purpose appears technically logical. The unintended consequence is that the strictest settings may effectively disable search engine access altogether, undermining the very purpose of having a discoverable website.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Practical implications: the SEO and traffic impact of blocked crawlers<\/h2>\n\n\n<p class=\"wp-block-paragraph\">For site operators who enable the AI Training block, the immediate consequence is that Googlebot and Bingbot are unable to retrieve critical indexing signals such as sitemaps. Without access to the sitemap, search engines cannot effectively discover new pages, understand site structure, or prioritize crawl frequency. Over time, this can lead to a decline in search visibility, reduced organic traffic, and a drop in rankings as search engines rely on stale or incomplete indexing.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Moreover, the error messages appearing in Google <a href=\"https:\/\/overcentral.com\/en\/search-console-indexing-report-freeze\/\" title=\"Google Search Console Stops Updating Indexing Report Since July 10\" data-iacss-internal=\"1\">Search Console<\/a> provide a clear indication of the impact. Search Console is an essential diagnostic tool for webmasters, and when it reports crawling errors caused by the AI Training block, the site owner is effectively receiving an official warning that their site&#8217;s search performance is being degraded. The longer the block remains active, the more severe the potential consequences become \u2014 particularly for high-traffic publishers who depend on search as their primary acquisition channel.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">It is also worth considering that the AI Training block applies to all pages, not just those containing advertising. Cloudflare&#8217;s statement about restricting access to pages with ads suggests a nuanced approach, but the Reddit report indicates that the block is being applied more broadly, at least in the case of the sitemap. This suggests that the implementation may be more aggressive than the policy description implies, or that the classification system requires further refinement to distinguish between legitimate search indexing and AI training activity.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">What are the potential causes of the Googlebot and Bingbot blocking?<\/h2>\n\n\n<p class=\"wp-block-paragraph\">The most likely explanation is that Cloudflare&#8217;s bot classification system groups Googlebot and Bingbot under a category that assumes AI training as a primary or secondary use case. Since these crawlers are owned by companies with robust AI initiatives, Cloudflare&#8217;s automated detection may be treating them as multi-purpose bots subject to the strictest restrictions. Another possibility is that the AI Training block uses a broad rule set that inadvertently applies to the IP ranges or user-agent strings associated with these search crawlers, especially if the crawlers are making requests that resemble AI scraping patterns. Additionally, the block may be activated by certain request patterns, such as high-frequency or deep-crawl behaviors that trigger AI-related mitigation rules designed to protect against aggressive data extraction.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Comparisons with other bot-blocking features and incidents<\/h2>\n\n\n<p class=\"wp-block-paragraph\">This is not the first time that security or bot management tools have interfered with search engine crawlers. Industry observers may recall previous incidents where CDN or firewall rules inadvertently blocked Googlebot due to overly aggressive rate limiting or misconfigured user-agent filters. In many of those cases, the issue was resolved through updates to the detection logic or through manual whitelisting of Google&#8217;s IP ranges.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What distinguishes the current situation is the policy-driven nature of the block. Unlike incidental rate-limiting errors, the AI Training block is an intentional feature with explicit user controls. This means that the issue is not a random bug but a structural consequence of how Cloudflare has chosen to implement its AI protections. Resolving it may require more than a simple patch; it may require Cloudflare to fundamentally reconsider how it classifies bots and under what circumstances the strictest settings are applied.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">There is also a strategic angle: Cloudflare is not alone in implementing AI-specific controls. Other providers and platforms have introduced similar restrictions, and the broader industry is grappling with how to balance content protection with open web discoverability. The incident with Cloudflare serves as a cautionary tale for other companies considering similar measures, illustrating that the lines between search bots and AI scrapers are increasingly blurred, and that any attempt to restrict one category may inadvertently restrict the other.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Best practices for website operators using Cloudflare AI bot settings<\/h2>\n\n\n<p class=\"wp-block-paragraph\">Given the current uncertainty, site owners who rely on Cloudflare must exercise caution when configuring the AI Training block. For those who prioritize search engine visibility, the safest approach is to disable the block until Cloudflare issues guidance or provides a separate setting specifically for AI training that does not affect Googlebot and Bingbot.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">However, for sites facing aggressive AI scraping, this is not an ideal solution. As a workaround, some operators may choose to use Cloudflare&#8217;s managed challenge or rate-limiting features as a less restrictive alternative, allowing legitimate search bots through while still imposing friction on known AI crawlers. Others may consider implementing robots.txt directives to explicitly instruct AI bots to avoid their content, though compliance with these directives is voluntary and not always guaranteed.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Another important step is to monitor Google Search Console closely for any errors or warnings related to sitemap access or crawl failures. If the AI Training block is activated, the Search Console will likely report these issues, providing clear evidence of the impact. Site operators should also check Bing Webmaster Tools for similar signals, as Bingbot is affected in the same manner.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Importantly, website owners should test any configuration changes in a staging environment before applying them to production sites, especially for high-traffic or mission-critical properties. The AI Training block should be considered a high-impact setting, and its effects on SEO should not be underestimated.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">When did Cloudflare announce the AI bot-blocking default change?<\/h2>\n\n\n<p class=\"wp-block-paragraph\">Cloudflare announced the forthcoming default change in early July, with the effective date set for September 15. The announcement clarified that from that date onward, multi-purpose crawlers\u2014including those used for both search and AI training\u2014would be subject to the strictest access restrictions unless webmasters manually adjusted their settings. The policy came amid growing global concern over unbridled data harvesting by AI companies and was positioned as a tool to empower website owners to protect their content.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Future outlook: the evolving relationship between CDNs, search engines, and AI<\/h2>\n\n\n<p class=\"wp-block-paragraph\">The Cloudflare incident illustrates a broader tension that will likely intensify in the coming years. As AI crawlers proliferate and become more sophisticated, the ability to distinguish between legitimate indexing and unauthorized scraping will become increasingly challenging. Search engines themselves are evolving into AI-driven entities, and the lines between search, discovery, and training are already starting to blur. Content delivery networks will need to adapt, developing more nuanced classification systems that can assess the intent and purpose of each bot request.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">At the same time, website owners face an increasingly complex decision matrix. On one hand, they want to protect their content from being harvested without compensation or acknowledgment. On the other hand, they rely on search engines to drive traffic and revenue. For many publishers, search is the lifeblood of their business, and any disruption to that flow is unacceptable. The Cloudflare incident highlights the challenges that arise when those two priorities conflict.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Looking forward, the industry may see the emergence of more granular consent frameworks that allow site owners to permit search indexing while explicitly blocking AI training, even when both activities are conducted by the same company. It is also possible that cloud providers will invest in machine learning-based detection systems that can infer the purpose of a crawler&#8217;s request from behavioral patterns, rather than relying on broad classifications based on user-agent strings or IP ranges.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For now, the incident serves as a reminder that the AI era is introducing new layers of complexity to web management, and that even well-intentioned features can have unintended consequences. Website operators should remain vigilant, monitor their analytics and Search Console reports, and stay informed about updates to Cloudflare&#8217;s policies and implementations. The situation also underscores the importance of community dialogue and shared knowledge, as demonstrated by the Reddit thread that brought this issue to broader attention.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Cloudflare has yet to respond, but given the volume of affected domains and the visibility of the issue, it is likely that an official statement will be forthcoming. Until then, site owners should proceed cautiously, prioritizing their search engine visibility while exploring less intrusive ways to manage AI scraping. The incident is a signal that the infrastructure of the web is shifting, and that site owners, CDNs, and search engines will need to work together to ensure that the web remains open, discoverable, and protected in equal measure.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>For website owners and SEO professionals who rely on Cloudflare\u2019s security and performance suite, a troubling discovery has surfaced: the platform\u2019s new AI bot-blocking feature appears to be inadvertently blocking the industry\u2019s most important web crawlers. Several site administrators have reported that enabling Cloudflare\u2019s &#8220;AI Training&#8221; block, a feature introduced to prevent artificial intelligence companies [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":75258,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/i.ibb.co\/23cHGHgr\/768310403-2175978963351511-6158792231091817705-n.webp","fifu_image_alt":"","footnotes":""},"categories":[31],"tags":[],"class_list":["post-75084","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/i.ibb.co\/23cHGHgr\/768310403-2175978963351511-6158792231091817705-n.webp","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75084","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=75084"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75084\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/75258"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=75084"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=75084"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=75084"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}