Court documents reveal OpenAI and Microsoft knew AI would destroy the web

Internal documents from the New York Times lawsuit reveal that OpenAI and Microsoft knew their AI would harm the web, calling it a 'doom loop'.

By Central
Court documents reveal OpenAI and Microsoft knew AI would destroy the web
Highlights
  • Microsoft's internal document admitted its AI content strategy started a 'doom loop' that would hurt model performance and the entire web.
  • Microsoft's Brent Hecht described the data scraping as 'the largest theft of labor in human history' in the unsealed court filings.
  • AI-driven summaries have already reduced referral traffic to news sites by as much as 60 percent, according to the court documents.

The unsealed court documents in the New York Times lawsuit against OpenAI and Microsoft have exposed a stunning admission: the companies knew their AI products would trigger a “doom loop” that would irreparably damage the web, characterized their data scraping as “the largest theft of labor in human history,” and conceded that their fair use defense made a “complete mockery” of the legal concept. These are not accusations from plaintiffs; they are the companies’ own internal warnings, now laid bare in 92 pages of filings that read less like legal briefing and more like a confession.

The Doom Loop Was Real: How Microsoft and OpenAI Anticipated the Collapse of Their Own Supply Chain

Perhaps the most damning piece of evidence comes from a Microsoft internal document that openly acknowledged the destructive trajectory of its AI content strategy. “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time,” the document states. “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.'”

This admission is striking not only for its candor but for its accuracy. As the court filing notes, Satya Nadella himself agreed under oath that conversing with chatbots “has substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source.” Microsoft later tried to soften the message, with spokesperson Alex Haurek telling The Verge that Nadella’s testimony “should not be confused with conclusions about copyright questions before the Court.” But the internal document leaves little room for interpretation: the company understood it was building a product that competes directly with the very publishers whose work it was ingesting.

The filing further quotes Microsoft admitting that “LLMs are a product that destroys its supply chain.” This is not a hypothetical — it is a self-fulfilling prophecy. The same data that trains the models is also the economic foundation of the organizations that produce it. By replacing the need to visit original sources, AI chatbots eliminate both revenue and incentive for publishers to continue creating the high-quality content that the models depend on. The doom loop is not a future risk; it is already underway.

What Did the Unsealed Court Documents Reveal About OpenAI and Microsoft?

The unsealed court documents reveal that OpenAI and Microsoft internally acknowledged that their AI models would cause significant harm to the web and the publishing industry. Microsoft’s Director of Applied Science, Brent Hecht, described the companies’ data scraping as “an astonishing theft of unprecedented proportions” and “the largest theft of labor in human history.” He also stated that Microsoft’s fair use defense would “make a complete mockery of the idea of ‘fair use.'” Additionally, internal documents acknowledged that the companies had started a “doom loop” that would damage both their models and the broader web, and that LLMs are “a product that destroys its supply chain.”

An Astonishing Theft: The Language of Confession in OpenAI and Microsoft’s Own Words

Much of the conversation around generative AI has focused on whether training on copyrighted material constitutes fair use or infringement. The newly unsealed documents show that at least some senior figures inside both companies had no doubt about the answer. Microsoft’s Brent Hecht did not mince words: the introduction of the filing directly quotes him as calling ChatGPT and Copilot’s harvesting of data “the largest theft of labor in human history” and saying that any defense under fair use “would arguably make a complete mockery of the idea of ‘fair use.'”

Microsoft tried to distance itself from Hecht’s statements. Alex Haurek said, “These comments reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.” In a separate filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterized Hecht as holding “divergent, academic, and forward-looking views” and stated that he is “not someone who speaks for Microsoft specifically as to his theoretical views on AI’s potential effect on content creators.”

But the document includes far more than Hecht’s opinions. OpenAI’s Head of ChatGPT (presumed to be Nick Turley) wrote that “publishers face an existential threat” from AI products, and that these products “are largely substitutive, period” and “will get more and more substitutive as they get better.” Far from being a single employee’s idiosyncratic view, the filing paints a picture of a company-wide awareness that the entire publishing business model was being undermined.

Paywall Circumvention and the Absence of Consent

The documents also reveal a troubling pattern regarding consent. While Satya Nadella is quoted as saying that “anything that is paywalled should be licensed,” the behavior on the ground told a different story. An OpenAI corporate representative admitted under testimony that he was “unaware of any effort to detect paywall content in its training datasets” or “to remove paywall content from its training datasets.” This discrepancy between public statements and internal practice underscores the gap between stated principles and operational reality.

The filing notes that “individuals within OpenAI and Microsoft ignored such issues as circumventing paywalls and violating terms of use when acquiring data.” The companies were not simply scraping publicly available content; they were systematically ingesting material behind paywalls, material that publishers relied on for their subscription revenue. This is not incidental — it is a direct attack on the revenue streams that sustain journalism.

Insanely Good at Regurgitation: OpenAI Knew Its Model Could Reproduce Copyrighted Work Verbatim

Another major revelation concerns OpenAI’s knowledge of its models’ ability to reproduce copyrighted content verbatim. As early as 2020, OpenAI recognized that its API “might output existing content verbatim.” By 2021, the company acknowledged that preventing memorization was important “for fair use [compliance] and minimizing copyright violations in model output.”

Yet by June 2022, internal communications show that employees understood that GPT-4 would have “memorized a ton of data and therefore will be insanely good at regurgitation.” This is not a bug — it is a feature of the technology that the companies chose to develop and deploy. The filing then provides multiple examples of ChatGPT outputting long strings of text taken directly from articles in the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to user queries. This is not transformative use; it is literal copying.

Substituting for Human Labor: The Modern Newsstand That Replaces Its Sources

OpenAI’s Policy Director Jack Clark wrote that “our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of the people that define the ‘culture’ of society.” Internal documents characterized ChatGPT as “the modern newsstand,” and employees bragged that the chatbot provides “fast, timely answers which you would have previously needed to go to a search engine for,” including “up-to-date sports scores, news, stock quotes, and more.”

Nick Turley, OpenAI’s Head of ChatGPT, went further, saying that once you get an answer from the chatbot, there is “no good reason to click” on a link to the source. This is a direct admission that the product is designed to keep users within the AI ecosystem, depriving publishers of the traffic they need to survive.

Referral Traffic Collapse: OpenAI’s Own Experts Blame AI Summaries

The filing also includes analysis from OpenAI’s own media and economic experts. Dr. Goldfarb, an OpenAI economic expert, concluded that “the decline in referral traffic to The Times’s properties has been driven by a combination” of factors including “Google AI Overviews.” Dr. Sinnreich, OpenAI’s media expert, noted that “referral traffic to publishers from both Google Search and Google Discover has dropped considerably — from over 5 billion monthly referrals via Discover to fewer than 4 billion, and from well over 3 billion via Search to slightly more than 2 billion — since Google introduced AI overviews.”

Both experts linked the precipitous drop directly to AI-generated summaries. Dr. Goldfarb specifically opined that “Google’s introduction of AI overviews may have depressed search referrals by 20 to 60 percent” for certain publishers. These are not external critics; these are experts hired by OpenAI who acknowledged that the very technology the company helped pioneer is destroying the traffic that publishers depend on.

Chasing Gazillions: The Profit Motive Behind the Doom Loop

Why would two of the world’s most sophisticated technology companies knowingly pursue a strategy that would decimate their own content supply chain? The answer may lie in a single word from OpenAI cofounder Greg Brockman. The filing notes that Brockman wrote he was “deeply motivated by the gazillions” he hoped to gain by commercializing OpenAI’s technology. The word “gazillions” is not a real number — it is a deliberate exaggeration that reveals a mindset driven by vast financial ambition.

OpenAI is reportedly planning an IPO based on a $1 trillion valuation. Microsoft has invested billions. The companies knew the damage they were causing but proceeded anyway because the potential rewards were too enormous to resist. As the filing states, “it’s clear that this came true” — the predictions of doom, the destruction of the web, the theft of labor — all of it happened exactly as the internal documents warned.

The Broader Implications: What This Means for the Future of the Web and Publishing

The unsealed documents represent a watershed moment in the ongoing legal battles between AI companies and content creators. They provide a rare window into the internal consciousness of the companies that have arguably done the most to reshape the internet in the last decade. The revelation that both OpenAI and Microsoft were fully aware of the destructive consequences of their actions — and pressed ahead anyway — undermines any claim of ignorance or good-faith innovation.

For publishers, the implications are existential. The data shows that AI-driven summaries have already slashed referral traffic by as much as 60 percent. If that trend continues, many news organizations may simply cease to exist, unable to sustain themselves on the remaining crumbs of direct traffic. The doom loop is not a theoretical concept; it is a present reality.

For regulators and lawmakers, the documents offer a clear case for action. If companies themselves admitted that they were engaging in “the largest theft of labor in human history” and that their products made a “mockery” of fair use, the argument for stronger copyright protections and clearer rules around AI training data becomes difficult to ignore. The European Union’s AI Act and ongoing lawsuits in the United States may gain new momentum from these admissions.

For the rest of us — the readers, the users, the consumers of both AI tools and traditional journalism — the documents serve as a stark reminder that convenience often comes at a hidden cost. Every time we ask a chatbot a question and receive a polished answer without clicking through to the source, we are participating in the slow dismantling of the information ecosystem that produced that knowledge in the first place. The doom loop is not just a corporate problem; it is one we are all complicit in.

The question now is whether the legal system will catch up to the technological reality, and whether the web can survive the model that was built on its ruins.

Share This Article