ChatGPT’s Search Index Serves Small Sites Too

Resoneo's analysis reveals OpenAI's labrador index treats partners and non-partners equally, giving small publishers visibility.

By Central
OpenAI's labrador index serves content from both licensed partners and independent sites without discrimination, democratizing search visibility.
Highlights
  • Resoneo analyzed 1,249 ChatGPT answers and found that the labrador index serves both partners and non-partners equally.
  • For free-tier users, the majority of ChatGPT search results come from OpenAI's own index rather than third-party search engines.
  • OpenAI has stopped tagging pipeline names, making it harder for publishers to verify how their content is retrieved.

For months, a central question has hovered over the relationship between OpenAI and the publishing world: Is the company’s search feature a closed garden for licensed partners, or does it offer genuine visibility to independent and smaller websites? New evidence from the French SEO consultancy Resoneo provides a definitive, data-driven answer. By analyzing 1,249 ChatGPT answers captured in July, Resoneo found that hundreds of outlets with no formal content deal with OpenAI were served by the company’s proprietary search index—codenamed “labrador”—in exactly the same manner as its licensed partners. For free-tier accounts, this in-house index handled the vast majority of ChatGPT search results, a finding that reshapes the strategic calculus for publishers evaluating whether to pursue a partnership or simply optimize their content for OpenAI’s crawler.

What the Labrador Index Actually Is

OpenAI’s server stream tags each web result returned to a user with the name of the pipeline that fetched it. Resoneo’s reverse-engineering revealed four distinct pipeline values, one of which was “labrador”—the company’s own index. Unlike the other pipelines that scrape results from third-party search engines like Google, labrador is built and maintained internally by OpenAI. Resoneo describes it as an index that is “topped up with press feeds and open science archives,” giving OpenAI direct, cost-free access to content without needing to pay a licensing fee to a third-party search provider.

The critical finding is that when comparing pages served from the labrador pipeline, the format, length, and freshness of the cited content were identical for both partners and non-partners. A licensing deal did not change how a page was stored, how much of it was displayed, or how recently its data was refreshed. This contradicts a prevailing assumption that OpenAI’s search capability was functionally a pay-to-play environment for established publishers only.

How the Measurement Was Conducted

Resoneo, a French SEO consultancy that sells optimization services and distributes a free Chrome extension for data capture, examined a corpus of 1,249 ChatGPT answers collected during July. The company classified each web result according to the pipeline tag attached by OpenAI’s server. This methodology provided a direct, server-level view of how content was being retrieved, rather than relying on inference from what appeared in a user’s browser.

In its free-account data, questions with settled answers—such as factual queries about history, science, or geography—local business searches, and product-related questions almost always returned results through the labrador index. News results were split more evenly, with roughly half coming from labrador and half scraped from Google. This suggests that OpenAI’s index is particularly robust for evergreen, structured, and commercially oriented content, while it still relies heavily on Google for the rapidly changing landscape of news.

The Data Behind the Paid and Free Tiers

The picture shifts notably when users move to paid accounts in “thinking” mode. Resoneo recorded 16,407 search results from this tier. Of those, approximately 75% were scraped from Google, while the in-house labrador index accounted for about 24%. This dramatic reversal indicates that OpenAI is actively investing more computational resources—or accessing broader data sources—for its paying customers, but it also means that for the vast majority of free users, the in-house index is the primary gateway.

This asymmetry has significant implications. For publishers optimizing for visibility, targeting the free tier—which constitutes the bulk of ChatGPT’s user base—means focusing on how their content performs within the labrador index. The paid tier’s heavier reliance on Google scraping introduces a separate set of dynamics, potentially tied to Google’s own ranking signals and content freshness.

A Correction That Reshaped the Narrative

Resoneo’s findings directly support a correction published in July by researcher Suganthan Mohanadasan. In June, Mohanadasan had described the labrador index as an “allowlist” of established publishers after examining ChatGPT’s network traffic. He reported that it “looks like a licensed tier,” naming domains such as Reuters, The Guardian, The Wall Street Journal, and Wikipedia as consistent presences. His initial interpretation suggested that smaller sites were effectively locked out of ChatGPT’s search results.

On July 14, Mohanadasan retracted that conclusion. A reader from Italy, using a free account, sent him captures showing that every publisher citation—including those from small Italian websites—went through the same labrador pipeline. Mohanadasan re-ran his own tests, acknowledged in his summary table that he had “over-reached” with the tier claim, and clarified that while the licensing deals are genuine, his earlier reading was based on viewing just one account’s perspective as representative of the entire system. “A single account only reveals how ChatGPT interacted with that one account,” he noted, recommending that future investigations use multiple accounts to get a clearer picture.

Why Did the Initial Reading Change?

The discrepancy between Mohanadasan’s first and second analyses highlights a fundamental challenge in reverse-engineering AI systems: pipeline tagging can vary by user, account type, geographic location, and time of testing. Resoneo’s dataset included both free and paid accounts, multiple countries, and logged-out sessions, with the same prompts replayed across different account types. This breadth provided a statistically more reliable view than Mohanadasan’s initial single-account approach, though his correction also drew on captures from two other readers’ accounts to validate the broader pattern.

Resoneo credits Mohanadasan’s work as the foundation for its own efforts, and the two investigations together form a coherent picture. Around July 21, however, OpenAI stopped tagging each search result with the name of the system that fetched it—the very tag both investigations had been reading. This change complicates future independent auditing of ChatGPT’s search behavior, making the July dataset a uniquely valuable snapshot.

What the Model Actually Sees from Your Page

Beyond the question of which index retrieves content, Resoneo’s analysis provides granular detail on how pages are stored and presented. The consultancy reviewed 534 pages that ChatGPT cited and compared each one with the snippets stored in OpenAI’s index. The results reveal a tightly constrained data structure.

Out of 463 pages that had an H1 heading, 387 snippets—or 83.6%—included it. The snippet is cut off just after 200 characters, typically drawn from the beginning of the page content rather than the meta description. (The Google-scrape pipeline still captures meta descriptions roughly one out of three times, but the in-house index prioritizes the actual page text.)

The median H1 heading was 51 characters long, which leaves approximately 150 characters for the body content that follows. This constrained space means that every element of a page’s template matters. A section kicker appears before the H1 on 29% of pages and consumes 18 characters. A publication date appears on 11% of pages, using 25 characters. The alt text of the first image appears on 9% of pages and can take up 50 characters on its own. In the sample, one out of every seven pages had no H1 markup at all. Resoneo notes that in those cases, the snippet begins with whatever subheading the template provides, which often results in a less coherent or less useful representation of the page’s core content.

Practical Implications for Publishers

For content creators and SEO professionals, these findings offer actionable intelligence. The labrador index stores a title and roughly 200 characters from the top of the page. Anything a template prints above the first paragraph—navigation bars, breadcrumbs, promotional banners, author bylines, dates, kickers—consumes part of that limited budget. The implication is clear: the opening structure of a page, including the H1 and the first 150 characters of body text, should be treated as a critical asset for AI retrievability.

Resoneo did not test whether changing a page’s H1 or opening content increases the likelihood of being cited. However, the data strongly suggests that pages with clear, concise, and well-structured openings are more likely to have their core message captured faithfully within that 200-character snippet. Pages that bury their thesis beneath template noise are at a disadvantage, regardless of whether they have a licensing deal.

Why the Licensing Debate Needs Reframing

The immediate takeaway for publishers is that a content deal with OpenAI is not a prerequisite for appearing in ChatGPT’s answers to free users. Resoneo’s evidence, supported by Mohanadasan’s correction, demonstrates that non-partner sites were already in the index that handles most of those answers. This does not mean licensing deals are worthless—they often involve financial compensation, prioritized crawling, feed-based ingestion, and potentially preferential placement in paid-tier results—but it does mean that visibility in the free tier is not a strong selling point for those agreements.

Whether a deal helps a page get cited more frequently is a different question, and neither investigation addressed it. Resoneo focused on how pages were stored and served, not on the statistical frequency of citation. Similarly, neither study examined whether partner articles appear more prominently, receive longer snippets, or are ranked higher within ChatGPT’s response generation.

What OpenAI’s Own Documentation Reveals

As of publication, OpenAI’s crawler page does not detail its in-house index or specify what its publisher agreements include. The company has not publicly confirmed the existence or scope of the labrador pipeline. Resoneo notes that partner articles reach OpenAI through a feed rather than a crawl, so a deal could change how content gets into the system—faster, more reliably, and possibly with richer metadata—even if the presentation of that content within an answer is identical to a non-partner page.

This opacity leaves publishers in a difficult position. They must decide whether to invest in formal partnerships without full transparency about what those partnerships deliver, or to optimize their pages for organic retrieval by OpenAI’s crawler without knowing the long-term stability of that access. The July data suggests that for the near future, the organic path is viable for most content types, particularly evergreen, local, and product-related information.

The Strategic Calculus for Independent Sites

The correction by Mohanadasan and the confirmation by Resoneo represent a significant shift in the perceived power dynamics of AI-driven search. Smaller sites, including niche publishers, local businesses, and independent blogs, should no longer assume they are invisible to ChatGPT. The labrador index appears to be a genuinely open retrieval system for free-tier users, provided the content meets relevance and quality thresholds.

This does not mean every small site will be equally represented. The index’s construction—topped up by press feeds and open science archives—suggests a bias toward certain types of authoritative or academic content. But the absence of a paywall or licensing gate means that the primary barrier to visibility is now technical and structural, rather than commercial. Publishers who invest in clean HTML, strong H1 structures, and concise opening paragraphs are likely to see better results from organic ChatGPT citations.

The July dataset also carries a warning: OpenAI’s decision to stop tagging pipeline names makes future independent auditing harder. Publishers cannot easily verify whether their content is being retrieved from labrador, scraped from Google, or ignored entirely. This loss of transparency shifts the burden onto publishers to monitor their own visibility through user reports and indirect measurements.

Ultimately, the question is no longer whether small sites can appear in ChatGPT’s search results—they can, and they do. The question is whether OpenAI will maintain this open architecture as it scales, monetizes, and negotiates with larger content partners. The evidence from July suggests a system that is currently more equitable than many assumed, but the dynamics of AI search are notoriously volatile. For now, the path to visibility runs through good content practices, not through a legal agreement.

Share This Article