AI Visibility Audit reveals why websites rank in search but vanish from ChatGPT

Discover why your website might be invisible to AI chatbots and how a 90-minute audit using free tools can fix it.

By Central
A five-step AI Visibility Audit from Common Crawl reveals the hidden configuration issues blocking sites from ChatGPT.
Highlights
  • A website's Google ranking does not guarantee visibility in generative AI systems like ChatGPT.
  • Common Crawl's CCBot determines whether a site's content enters the training data for major AI models.
  • Most AI invisibility issues stem from simple configuration errors in robots.txt, CDN settings, or firewall rules.

A website can dominate Google’s search results and still be completely invisible to ChatGPT, Gemini, Claude, or Perplexity. The gap between classical search visibility and AI visibility is growing, and the root causes often have nothing to do with traditional SEO. An AI Visibility Audit reveals exactly why a site vanishes from generative AI systems while ranking at the top of Google — and the answer usually lies in configuration issues that many website owners do not even know exist.

Stephen Burns, Web Intelligence Lead at the Common Crawl Foundation, published a practical handbook for SEOs and GEOs outlining a structured five-step audit that diagnoses AI invisibility. The entire process takes approximately 90 minutes, uses exclusively free tools, and addresses a fundamental blind spot in modern digital strategy. Before any on-page optimization, technical SEO, or link building matters, a website must first be reachable by the crawlers that collect training data for large language models. If that pipeline is blocked, the site simply does not exist in the model’s knowledge base.

Why Google rankings do not guarantee AI visibility

The core insight is counterintuitive but well-documented. Google’s indexing pipeline and the data pipelines feeding AI models are entirely separate systems. A site can be fully optimized for Google’s algorithms, with strong technical SEO, high-quality content, and authoritative backlinks, yet remain absent from ChatGPT’s training corpus or Perplexity’s live retrieval index. The reason is often a handful of lines in a robots.txt file, a default setting in a content delivery network, or a firewall rule that blocks AI crawlers without the site owner’s knowledge.

Common Crawl, the nonprofit organization founded by Gil Elbaz in 2007, operates a bot called CCBot that crawls the open web monthly and deposits the results as freely downloadable archives on Amazon S3. These archives became one of the earliest and most significant data sources for training models from OpenAI, Anthropic, Google, and others. Whether CCBot is permitted to access a website effectively determines whether that site’s content enters the training data that shapes what language models know and how they respond.

Since 2008, Common Crawl has accumulated over 10 petabytes of data. The crawler captures approximately 2 to 2.5 billion web pages per month, with a total archive exceeding 300 billion pages. A single monthly crawl contains roughly 350 to 400 terabytes of uncompressed data. A website that is not in that crawl is not in the training data, regardless of its Google ranking.

The four steps from the open web to the model

Understanding how content moves from a website into an AI model clarifies where the blockages occur. First, CCBot fetches publicly accessible pages, follows links, and respects robots.txt directives. Sites that are not crawled at this stage are simply absent from the crawl output. Second, AI providers access the Common Crawl archive, filter for quality, and use the data to train their models. Third, the trained model internalizes the content in its parametric memory. Fourth, when a user queries the model, it can retrieve and reformulate that information.

The process is sequential and each step introduces delay. A website may go live today, but the CCBot may not discover it until the next crawl cycle. The monthly crawl archive containing that site must then be published, and only a subsequent training round by the AI provider will incorporate the data. Visibility in AI systems builds slowly, and the timing is dictated by the model’s training schedule, not the website owner’s publishing calendar.

Harmonic Centrality versus PageRank

Common Crawl publishes a Web Graph that maps the link structure of the web at the host and domain level. From this data, the organization derives a metric called Harmonic Centrality. The concept is straightforward: a domain located near the core of the link network is considered more central than one at the periphery, and centrally positioned domains receive preferential treatment during crawling.

This distinguishes Harmonic Centrality from classical PageRank. PageRank measures how many important pages link to a given page. Harmonic Centrality measures how close a domain sits to the structural center of the web. A single link from a highly central page can boost centrality more effectively than dozens of links from isolated, peripheral pages. For practitioners, this means link quality has a double payoff. Higher centrality translates to more frequent crawls, more pages captured, more training material ingested, and a greater likelihood that models know and recommend the site.

Two layers of AI visibility: parametric memory and live retrieval

Once a site is captured in training data, it exists in two distinct layers within an AI system. The first is parametric memory: content captured before the training cutoff date and embedded in the model’s fixed weights. This is the knowledge the model retrieves internally without any external lookup. The second layer is live retrieval, often implemented through Retrieval-Augmented Generation. In this mode, the model fetches fresh content at query time from live, crawlable sources.

A website may be present in parametric memory but blocked for live retrieval, or accessible for live retrieval but absent from the training corpus. A complete audit must examine both dimensions. Was the site reachable when the model was trained, and is it accessible today for live queries? This distinction, formalized in a framework by Duane Forrester, becomes critical as models age. The further a model’s training cutoff date recedes into the past, the more it relies on live retrieval for current topics, and the more the site’s present-day crawler accessibility determines whether it appears in responses.

Each model behaves differently. ChatGPT, Gemini, and Claude blend trained knowledge with live search results. Perplexity operates predominantly through live retrieval. The specific cutoff dates change with each model version, and only the providers themselves publish authoritative information about their training schedules.

Why most sites are blocked without realizing it

The most common reason for AI invisibility is a default configuration in a CDN or firewall that no one has reviewed. Two mechanisms account for the vast majority of cases. Some CDNs automatically write disallow rules for AI crawlers into the robots.txt file, blocking GPTBot, ClaudeBot, Google-Extended, and CCBot by default. Other blocks are enforced at the web application firewall level, where a toggle labeled “Block AI bots” rejects requests before they reach the origin server. The site owner sees nothing in their own server logs, and the bot receives a 403 response.

Because a single large CDN sits in front of a substantial portion of the web, and because some providers default to blocking AI crawlers for new domains, this situation is closer to the norm than the exception. The impact on news and publishing is particularly severe. A study of 100 major US and UK news sites conducted in January 2026 found that CCBot was the most restricted training bot, blocked by 75 percent of sites, followed by Anthropic-ai at 72 percent, ClaudeBot at 69 percent, and GPTBot at 62 percent. For any publisher or news client, checking CCBot access should be the first priority.

Selected data from the study illustrates the trend:

  • arXiv study of news sites: AI bots blocked on 23 percent of sites in September 2023, rising to 60 percent by May 2025
  • Originality.AI study of top 1,000 sites: GPTBot blocked on 35.7 percent of sites as of August 2024
  • Originality.AI study of top 1,000 sites: CCBot blocked on 22.1 percent of sites
  • BuzzStream study of 100 news sites: CCBot blocked on 75 percent of sites as of January 2026

The remedy is often as simple as the cause. Because crawlers like CCBot respect robots.txt, restoring access typically requires nothing more than allowing the relevant AI agents in the robots.txt file and at the CDN or firewall level. Once access is reopened, the crawlers resume their normal rhythm. The earlier a site opens its doors, the sooner it begins building presence in the monthly archives, and that presence compounds over time. Opening access today is the single most effective and simplest step most websites can take toward AI visibility.

The legitimate choice to opt out

Not every publisher wants to be included in AI training. Some choose exclusion on principle, for economic reasons, or out of objection to how their content is used. That is a legitimate decision. However, the opt-out mechanisms are less reliable than many assume. Different crawlers respond to different directives, and a rule that stops one bot may leave another unaffected. The same access checks prescribed in the AI Visibility Audit are equally useful for verifying that an opt-out is actually working as intended.

English bias in AI systems: what it means for non-English sites

For websites operating in languages other than English, the challenge is compounded. A site may rank well in local search results yet lose AI citations to English-language competitors. The root cause is corpus composition. English accounts for approximately 43 to 45 percent of pages in the current Common Crawl index, and after quality filtering, the effective proportion is even higher. This drives two distinct biases.

The retrievers used in AI search systems tend to favor English and authoritative domains over comparable localized sites, particularly when the query is in English. Additionally, large models appear to operate in a partially language-agnostic conceptual space that tilts toward English representations. The practical lever is the retriever: a high-quality English version of the most important pages expands reach. Clean hreflang annotations, dedicated crawlable URLs, and genuinely well-written English content — not machine translation — all help. Both language versions must pass the same access checks.

It is important to be precise when communicating this to clients. Models tend toward English, but they do not translate into English. Anthropic, for example, describes Claude as operating on shared, language-independent concepts rather than performing linguistic translation.

The five steps of an AI Visibility Audit

The audit follows a sequential process because an early problem often explains issues discovered later. Every step uses free tools.

Step 1: Can the crawler reach the site?

This step checks whether anything blocks specific AI crawlers. First, inspect the robots.txt for disallow directives targeting CCBot, GPTBot, ClaudeBot, and Google-Extended. A clean robots.txt means nothing if the firewall still rejects the bot. Therefore, query the server using the CCBot user-agent string and compare the response to that of a standard browser. A 200 status for CCBot indicates access is open. A 403 status confirms the site is blocked. If the browser returns 200 and CCBot returns 403, the block is at the CDN or firewall level. Resolution involves checking the CDN configuration and the web server settings. Some hosting providers block AI bots by default, which may require contacting support or switching providers.

Step 2: Is the domain actually in the archive?

Use the Common Crawl Index Server to check whether the domain appears in the archive, when it was last crawled, and approximately how many pages were captured. No results indicate the domain is not included. Very few results suggest only superficial coverage. Old timestamps mean the data is stale. A domain can be technically accessible yet still receive very few crawler visits.

Intermediary step: Is it actually the real CCBot?

The genuine CCBot operates from fixed IP ranges with reverse DNS entries. When logging a bot’s IP, verify it using forward-confirmed reverse DNS. The real CCBot resolves to a hostname under .crawl.commoncrawl.org, which in turn resolves back to the same IP. Impersonators fail this check.

Step 3: Is the domain prioritized or deprioritized?

Check the domain’s Harmonic Centrality and its rank within the Web Graph. A low rank means the domain is deprioritized in the crawl budget. Even with open access, the crawler will visit such a site infrequently and only superficially. A low centrality score should be flagged as a strategic risk and targeted for link building aimed at central, well-connected domains. The fastest current tool for this check is the CC Rank Checker at webgraph.metehan.ai, built by Metehan Yesilyurt using Web Graph data. Common Crawl is developing its own tool.

Step 4: Can the content be cleanly represented?

Entities without structured data are harder for models to represent in training and more difficult for models to associate correctly. On important pages, verify the Schema.org markup for Organization, Article or Product, Author, and Breadcrumb. The Google Rich Results Test provides a straightforward way to check.

Step 5: Does the content exist without JavaScript?

Many AI crawlers behave like early Googlebot: they fetch the raw HTML but do not execute JavaScript. If critical content is loaded only after JavaScript execution, the crawler may capture an empty shell. Compare the raw server response with the rendered page. Search the raw HTML for a key text fragment from the page. A match confirms server-side rendering and crawler visibility. No match indicates the content is loaded via JavaScript and may be invisible to the crawler. Testing with JavaScript disabled in the browser or using a tool like Screaming Frog can help confirm.

Tools and output

The primary tools for the audit are the Common Crawl Index Server for archive presence and last-crawl timestamps, and the CC Rank Checker for centrality and crawl priority. Four additional utilities support the process: Screaming Frog for custom user-agent crawls, curl for rapid firewall and raw HTML tests, the Google Rich Results Test for markup validation, and the CDN dashboard where most accidental blocks are actually resolved.

The findings should be compiled per domain on a single-page scorecard. Each check receives a pass or fail status for access and rendering, a confirmed presence or absence in the index, a centrality grade, and an inventory of missing structured data, each accompanied by a specific remediation action.

AI visibility is not a substitute for classical SEO, but it is rapidly becoming a distinct and necessary discipline. The gap between Google rankings and ChatGPT visibility is real, measurable, and usually caused by solvable configuration issues. An audit that takes ninety minutes and uses only free tools can determine whether a site is truly present in the AI ecosystem or effectively invisible. For most websites, the answer is not a matter of content quality or backlinks. It is a matter of access — and access is the easiest problem to fix.

Share This Article