{"id":75648,"date":"2026-08-12T00:26:34","date_gmt":"2026-08-12T04:26:34","guid":{"rendered":"https:\/\/overcentral.com\/en\/?p=75648"},"modified":"2026-08-12T00:26:34","modified_gmt":"2026-08-12T04:26:34","slug":"chatgpt-search-index-small-sites","status":"publish","type":"post","link":"https:\/\/overcentral.com\/en\/chatgpt-search-index-small-sites\/","title":{"rendered":"ChatGPT&#8217;s Search Index Serves Small Sites Too"},"content":{"rendered":"<p>For months, a central question has hovered over the relationship between OpenAI and the publishing world: Is the company\u2019s search feature a closed garden for licensed partners, or does it offer genuine visibility to independent and smaller websites? New evidence from the French SEO consultancy Resoneo provides a definitive, data-driven answer. By analyzing 1,249 ChatGPT answers captured in July, Resoneo found that hundreds of outlets with no formal content deal with OpenAI were served by the company\u2019s proprietary search index\u2014codenamed \u201clabrador\u201d\u2014in exactly the same manner as its licensed partners. For free-tier accounts, this in-house index handled the vast majority of ChatGPT <a href=\"https:\/\/overcentral.com\/en\/claude-ai-shared-chats-leak\/\" title=\"Claude AI Shared Chats Leak into Google Search Results\" data-iacss-internal=\"1\">search results<\/a>, a finding that reshapes the strategic calculus for publishers evaluating whether to pursue a partnership or simply optimize their content for OpenAI\u2019s crawler.<\/p>\n<h2>What the Labrador Index Actually Is<\/h2>\n<p>OpenAI\u2019s server stream tags each web result returned to a user with the name of the pipeline that fetched it. Resoneo\u2019s reverse-engineering revealed four distinct pipeline values, one of which was \u201clabrador\u201d\u2014the company\u2019s own index. Unlike the other pipelines that scrape results from third-party search engines like <a href=\"https:\/\/www.google.com\/\" target=\"_blank\" rel=\"noopener noreferrer\" data-iacss-external=\"1\">Google<\/a>, labrador is built and maintained internally by OpenAI. Resoneo describes it as an index that is \u201ctopped up with press feeds and open science archives,\u201d giving OpenAI direct, cost-free access to content without needing to pay a licensing fee to a third-party search provider.<\/p>\n<p>The critical finding is that when comparing pages served from the labrador pipeline, the format, length, and freshness of the cited content were identical for both partners and non-partners. A licensing deal did not change how a page was stored, how much of it was displayed, or how recently its data was refreshed. This contradicts a prevailing assumption that OpenAI\u2019s search capability was functionally a pay-to-play environment for established publishers only.<\/p>\n<h2>How the Measurement Was Conducted<\/h2>\n<p>Resoneo, a French SEO consultancy that sells optimization services and distributes a free Chrome extension for data capture, examined a corpus of 1,249 ChatGPT answers collected during July. The company classified each web result according to the pipeline tag attached by OpenAI\u2019s server. This methodology provided a direct, server-level view of how content was being retrieved, rather than relying on inference from what appeared in a user\u2019s browser.<\/p>\n<p>In its free-account data, questions with settled answers\u2014such as factual queries about history, science, or geography\u2014local business searches, and product-related questions almost always returned results through the labrador index. News results were split more evenly, with roughly half coming from labrador and half scraped from Google. This suggests that OpenAI\u2019s index is particularly robust for evergreen, structured, and commercially oriented content, while it still relies heavily on Google for the rapidly changing landscape of news.<\/p>\n<h2>The Data Behind the Paid and Free Tiers<\/h2>\n<p>The picture shifts notably when users move to paid accounts in \u201cthinking\u201d mode. Resoneo recorded 16,407 search results from this tier. Of those, approximately 75% were scraped from Google, while the in-house labrador index accounted for about 24%. This dramatic reversal indicates that OpenAI is actively investing more computational resources\u2014or accessing broader data sources\u2014for its paying customers, but it also means that for the vast majority of free users, the in-house index is the primary gateway.<\/p>\n<p>This asymmetry has significant implications. For publishers optimizing for visibility, targeting the free tier\u2014which constitutes the bulk of ChatGPT\u2019s user base\u2014means focusing on how their content performs within the labrador index. The paid tier\u2019s heavier reliance on Google scraping introduces a separate set of dynamics, potentially tied to Google\u2019s own ranking signals and content freshness.<\/p>\n<h2>A Correction That Reshaped the Narrative<\/h2>\n<p>Resoneo\u2019s findings directly support a correction published in July by researcher Suganthan Mohanadasan. In June, Mohanadasan had described the labrador index as an \u201callowlist\u201d of established publishers after examining ChatGPT\u2019s network traffic. He reported that it \u201clooks like a licensed tier,\u201d naming domains such as Reuters, The Guardian, The Wall Street Journal, and Wikipedia as consistent presences. His initial interpretation suggested that smaller sites were effectively locked out of ChatGPT\u2019s search results.<\/p>\n<p>On July 14, Mohanadasan retracted that conclusion. A reader from Italy, using a free account, sent him captures showing that every publisher citation\u2014including those from small Italian websites\u2014went through the same labrador pipeline. Mohanadasan re-ran his own tests, acknowledged in his summary table that he had \u201cover-reached\u201d with the tier claim, and clarified that while the licensing deals are genuine, his earlier reading was based on viewing just one account\u2019s perspective as representative of the entire system. \u201cA single account only reveals how ChatGPT interacted with that one account,\u201d he noted, recommending that future investigations use multiple accounts to get a clearer picture.<\/p>\n<h2>Why Did the Initial Reading Change?<\/h2>\n<p>The discrepancy between Mohanadasan\u2019s first and second analyses highlights a fundamental challenge in reverse-engineering AI systems: pipeline tagging can vary by user, account type, geographic location, and time of testing. Resoneo\u2019s dataset included both free and paid accounts, multiple countries, and logged-out sessions, with the same prompts replayed across different account types. This breadth provided a statistically more reliable view than Mohanadasan\u2019s initial single-account approach, though his correction also drew on captures from two other readers\u2019 accounts to validate the broader pattern.<\/p>\n<p>Resoneo credits Mohanadasan\u2019s work as the foundation for its own efforts, and the two investigations together form a coherent picture. Around July 21, however, OpenAI stopped tagging each search result with the name of the system that fetched it\u2014the very tag both investigations had been reading. This change complicates future independent auditing of ChatGPT\u2019s search behavior, making the July dataset a uniquely valuable snapshot.<\/p>\n<h2>What the Model Actually Sees from Your Page<\/h2>\n<p>Beyond the question of which index retrieves content, Resoneo\u2019s analysis provides granular detail on how pages are stored and presented. The consultancy reviewed 534 pages that ChatGPT cited and compared each one with the snippets stored in OpenAI\u2019s index. The results reveal a tightly constrained data structure.<\/p>\n<p>Out of 463 pages that had an H1 heading, 387 snippets\u2014or 83.6%\u2014included it. The snippet is cut off just after 200 characters, typically drawn from the beginning of the page content rather than the meta description. (The Google-scrape pipeline still captures meta descriptions roughly one out of three times, but the in-house index prioritizes the actual page text.)<\/p>\n<p>The median H1 heading was 51 characters long, which leaves approximately 150 characters for the body content that follows. This constrained space means that every element of a page\u2019s template matters. A section kicker appears before the H1 on 29% of pages and consumes 18 characters. A publication date appears on 11% of pages, using 25 characters. The alt text of the first image appears on 9% of pages and can take up 50 characters on its own. In the sample, one out of every seven pages had no H1 markup at all. Resoneo notes that in those cases, the snippet begins with whatever subheading the template provides, which often results in a less coherent or less useful representation of the page\u2019s core content.<\/p>\n<h2>Practical Implications for Publishers<\/h2>\n<p>For content creators and SEO professionals, these findings offer actionable intelligence. The labrador index stores a title and roughly 200 characters from the top of the page. Anything a template prints above the first paragraph\u2014navigation bars, breadcrumbs, promotional banners, author bylines, dates, kickers\u2014consumes part of that limited budget. The implication is clear: the opening structure of a page, including the H1 and the first 150 characters of body text, should be treated as a critical asset for AI retrievability.<\/p>\n<p>Resoneo did not test whether changing a page\u2019s H1 or opening content increases the likelihood of being cited. However, the data strongly suggests that pages with clear, concise, and well-structured openings are more likely to have their core message captured faithfully within that 200-character snippet. Pages that bury their thesis beneath template noise are at a disadvantage, regardless of whether they have a licensing deal.<\/p>\n<h2>Why the Licensing Debate Needs Reframing<\/h2>\n<p>The immediate takeaway for publishers is that a content deal with OpenAI is not a prerequisite for appearing in ChatGPT\u2019s answers to free users. Resoneo\u2019s evidence, supported by Mohanadasan\u2019s correction, demonstrates that non-partner sites were already in the index that handles most of those answers. This does not mean licensing deals are worthless\u2014they often involve financial compensation, prioritized crawling, feed-based ingestion, and potentially preferential placement in paid-tier results\u2014but it does mean that visibility in the free tier is not a strong selling point for those agreements.<\/p>\n<p>Whether a deal helps a page get cited more frequently is a different question, and neither investigation addressed it. Resoneo focused on how pages were stored and served, not on the statistical frequency of citation. Similarly, neither study examined whether partner articles appear more prominently, receive longer snippets, or are ranked higher within ChatGPT\u2019s response generation.<\/p>\n<h2>What OpenAI\u2019s Own Documentation Reveals<\/h2>\n<p>As of publication, OpenAI\u2019s crawler page does not detail its in-house index or specify what its publisher agreements include. The company has not publicly confirmed the existence or scope of the labrador pipeline. Resoneo notes that partner articles reach OpenAI through a feed rather than a crawl, so a deal could change how content gets into the system\u2014faster, more reliably, and possibly with richer metadata\u2014even if the presentation of that content within an answer is identical to a non-partner page.<\/p>\n<p>This opacity leaves publishers in a difficult position. They must decide whether to invest in formal partnerships without full transparency about what those partnerships deliver, or to optimize their pages for organic retrieval by OpenAI\u2019s crawler without knowing the long-term stability of that access. The July data suggests that for the near future, the organic path is viable for most content types, particularly evergreen, local, and product-related information.<\/p>\n<h2>The Strategic Calculus for Independent Sites<\/h2>\n<p>The correction by Mohanadasan and the confirmation by Resoneo represent a significant shift in the perceived power dynamics of AI-driven search. Smaller sites, including niche publishers, local businesses, and independent blogs, should no longer assume they are invisible to ChatGPT. The labrador index appears to be a genuinely open retrieval system for free-tier users, provided the content meets relevance and quality thresholds.<\/p>\n<p>This does not mean every small site will be equally represented. The index\u2019s construction\u2014topped up by press feeds and open science archives\u2014suggests a bias toward certain types of authoritative or academic content. But the absence of a paywall or licensing gate means that the primary barrier to visibility is now technical and structural, rather than commercial. Publishers who invest in clean HTML, strong H1 structures, and concise opening paragraphs are likely to see better results from organic ChatGPT citations.<\/p>\n<p>The July dataset also carries a warning: OpenAI\u2019s decision to stop tagging pipeline names makes future independent auditing harder. Publishers cannot easily verify whether their content is being retrieved from labrador, scraped from Google, or ignored entirely. This loss of transparency shifts the burden onto publishers to monitor their own visibility through user reports and indirect measurements.<\/p>\n<p>Ultimately, the question is no longer whether small sites can appear in ChatGPT\u2019s search results\u2014they can, and they do. The question is whether OpenAI will maintain this open architecture as it scales, monetizes, and negotiates with larger content partners. The evidence from July suggests a system that is currently more equitable than many assumed, but the dynamics of <a href=\"https:\/\/overcentral.com\/en\/ai-search-fresh-content\/\" title=\"AI Search Rewards Fresh Content and Community Validation\" data-iacss-internal=\"1\">AI search<\/a> are notoriously volatile. For now, the path to visibility runs through good content practices, not through a legal agreement.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>For months, a central question has hovered over the relationship between OpenAI and the publishing world: Is the company\u2019s search feature a closed garden for licensed partners, or does it offer genuine visibility to independent and smaller websites? New evidence from the French SEO consultancy Resoneo provides a definitive, data-driven answer. By analyzing 1,249 ChatGPT [&hellip;]<\/p>\n","protected":false},"author":7,"featured_media":75652,"comment_status":"closed","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"fifu_image_url":"https:\/\/raw.githubusercontent.com\/medeiroslima\/overcentral-images\/main\/images\/ocie_1786508810091.jpg","fifu_image_alt":"ChatGPT's Search Index Serves Small Sites Too","footnotes":""},"categories":[31],"tags":[],"class_list":["post-75648","post","type-post","status-publish","format-standard","has-post-thumbnail","category-technology"],"fifu_image_url":"https:\/\/raw.githubusercontent.com\/medeiroslima\/overcentral-images\/main\/images\/ocie_1786508810091.jpg","fifu_image_alt":"ChatGPT's Search Index Serves Small Sites Too","_links":{"self":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75648","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/users\/7"}],"replies":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/comments?post=75648"}],"version-history":[{"count":0,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/posts\/75648\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media\/75652"}],"wp:attachment":[{"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/media?parent=75648"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/categories?post=75648"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/overcentral.com\/en\/wp-json\/wp\/v2\/tags?post=75648"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}