Seattle Times, Newsday launch lawsuit against OpenAI, Microsoft

Two major newspapers allege that OpenAI and Microsoft built AI products using stolen journalism, threatening the news industry.

By Central
The lawsuit argues that generative AI functions as a 'rapacious consumer' of original journalism, undermining its economic foundations.
Highlights
  • The Seattle Times and Newsday filed a lawsuit against OpenAI and Microsoft for copyright infringement.
  • The lawsuit claims that ChatGPT and Copilot were trained on millions of articles without authorization.
  • The complaint warns that unchecked AI could render the news industry 'broken beyond repair'.

The Seattle Times and Newsday have filed a lawsuit against OpenAI and Microsoft, alleging that the companies engaged in widespread copyright infringement by using the newspapers’ journalism to train their artificial intelligence products. The complaint, lodged in the U.S. District Court for the Southern District of New York, argues that the rapid advancement of generative AI threatens to dismantle the economic foundations of the news industry, rendering it “broken beyond repair.” This legal action marks a significant escalation in the ongoing conflict between publishers and the technology companies whose AI systems depend on vast quantities of human-authored text to function.

Two Major Newspapers Allege AI Giants Built Products on Stolen Journalism

The lawsuit contends that OpenAI and Microsoft built their flagship AI products, including ChatGPT and the Copilot assistant, by systematically copying and processing millions of articles from The Seattle Times and Newsday without authorization or compensation. The publishers argue that these AI models do not simply learn from text in a transformative way but instead function as “rapacious consumers” of original journalism, ultimately producing outputs that mimic, summarize, or rephrase the very content they ingested. The legal filing warns that without intervention, the entire ecosystem of professional journalism could collapse under the weight of a technology that feeds on its work.

The complaint draws a vivid parallel, describing generative AI as “a snake eating its own tail.” This metaphor captures a central paradox: AI systems touted as engines of creation are, in the view of the plaintiffs, fundamentally dependent on the uncompensated extraction of human creativity. If news organizations cannot protect their intellectual property, the lawsuit argues, they will lack the economic incentive to produce the original reporting that AI systems require to remain useful. The result, the publishers claim, would be a self-defeating cycle in which the quality of AI outputs degrades as the source material dries up.

What Did the Seattle Times and Newsday Allege in Their Lawsuit Against OpenAI and Microsoft?

The Seattle Times and Newsday alleged that OpenAI and Microsoft engaged in the unauthorized reproduction of their copyrighted articles to train large language models powering ChatGPT and Copilot. The lawsuit describes these AI products as engines of “copying and derivative imitation” rather than genuine creation, arguing that the companies benefited commercially from the uncompensated use of the publishers’ journalism. The plaintiffs further contend that this practice constitutes a direct threat to the survival of professional news organizations, as it undermines their ability to monetize their work and sustain investigative reporting.

The action by The Seattle Times and Newsday does not occur in a vacuum. In December 2023, The New York Times filed a landmark copyright infringement lawsuit against OpenAI and Microsoft, setting a precedent that has since spurred a wave of similar litigation from authors, visual artists, and music publishers. The New York Times case, which remains ongoing, alleged that OpenAI’s training data included millions of the newspaper’s articles and that ChatGPT could reproduce substantial portions of that content verbatim. The Seattle Times and Newsday lawsuit follows the same legal theory, arguing that the defendants engaged in willful infringement on a massive scale.

What distinguishes this latest lawsuit is the sheer breadth of the plaintiffs’ concerns. The complaint does not merely seek financial damages; it asks the court to order the destruction of AI models that were trained on the publishers’ content. Such a remedy, if granted, would represent a seismic disruption to OpenAI and Microsoft’s operations, potentially invalidating years of development and billions of dollars in investment. The publishers are also seeking injunctive relief to prevent further use of their journalism in AI training without explicit licensing agreements.

The legal strategy of the news organizations is becoming increasingly coordinated. Several digital news outlets, including The Intercept and Raw Story, have filed similar lawsuits, and a growing number of publishers are exploring collective action. The Seattle Times and Newsday, however, represent a particularly important constituency: regional newspapers that serve specific communities and rely on subscription and advertising revenue to fund local journalism. If these publishers cannot protect their content, the argument goes, the public interest in local news coverage will suffer disproportionately.

The Irony of Funding: Microsoft and OpenAI Supported Seattle Times Journalism

A notable and awkward dimension of the Seattle Times lawsuit is the preexisting relationship between the newspaper and the technology companies it is now suing. Microsoft and OpenAI have, in the past, provided financial support for some of the newspaper’s journalism projects and fellowships. This funding, which the lawsuit acknowledges, was framed as a philanthropic investment in independent journalism. The Microsoft spokesperson, responding to the lawsuit, expressed surprise at the legal action, telling GeekWire that the company is “always happy to sit down and explore solutions to this type of dispute.”

This dynamic adds a layer of complexity to the case. The plaintiffs are not arm’s-length adversaries with no prior connection to the defendants. They are organizations that accepted financial support from the very companies they now accuse of systematically undermining their business model. The lawsuit, however, draws a clear line between the charitable funding of specific projects and the wholesale appropriation of the newspaper’s entire archive for commercial AI training. The publishers argue that accepting grants for a fellowship program does not constitute a waiver of copyright protection for the broader body of their work.

For Microsoft, the situation is particularly delicate. The company has positioned itself as a champion of responsible AI development and has made public commitments to support journalism through licensing agreements and other partnerships. The lawsuit threatens to undermine that narrative by portraying Microsoft as a company that funds journalism with one hand while cannibalizing its economic base with the other. The Microsoft spokesperson’s statement, which emphasized a willingness to “explore solutions,” suggests that the company is eager to avoid a protracted legal battle that could generate damaging headlines and set unfavorable precedents.

How Generative AI Consumes and Reproduces Copyrighted Journalism

To understand the legal stakes of the lawsuit, it is necessary to examine the technical process by which AI models like ChatGPT and Copilot are created. Large language models are trained on vast datasets that include billions of words scraped from the public internet, including articles from news websites, blogs, books, and academic journals. The training process involves analyzing the statistical relationships between words and phrases, allowing the model to generate coherent text that mimics the patterns it has learned.

The critical issue in the lawsuit is not whether the models learn from the text but whether that learning constitutes copyright infringement. OpenAI and Microsoft have argued that training AI on publicly available text is a form of fair use, analogous to a human reading a book and learning from its content. The publishers, however, contend that the scale and nature of the copying are fundamentally different. When a human reads a newspaper article, they do not create a permanent digital copy of the entire text in a database that can be queried for commercial purposes. An AI training process, by contrast, involves the systematic reproduction of the original work in its entirety, often multiple times, as part of the model’s training pipeline.

Furthermore, the publishers argue that the outputs of these AI models can infringe on their copyrights by reproducing substantial portions of the original articles. The New York Times lawsuit famously demonstrated that ChatGPT could generate near-verbatim reproductions of its articles when prompted in specific ways. The Seattle Times and Newsday lawsuit makes similar allegations, arguing that the models are capable of producing “copies and derivative imitations” of the journalism they consumed. This capacity for reproduction, the plaintiffs argue, directly competes with the original publishers by providing users with access to the information without visiting the newspaper’s website or paying for a subscription.

What Is the “Snake Eating Its Own Tail” Argument in the OpenAI Lawsuit?

The “snake eating its own tail” metaphor used in the Seattle Times and Newsday lawsuit describes a self-destructive feedback loop in which generative AI consumes the journalism it was trained on, thereby undermining the economic viability of the organizations that produce that journalism. As AI systems become more capable of generating news-like content, users may rely less on original news sources, reducing traffic and revenue for publishers. With less revenue, publishers are forced to cut back on reporting, reducing the volume and quality of new journalism available for future AI training. Over time, the AI models degrade because they are trained on an increasingly thin and recycled dataset, while the news industry itself collapses. The metaphor captures the existential risk that the plaintiffs believe AI poses not just to their own businesses but to the entire information ecosystem.

The Broader War Between Publishers and AI Platforms

The Seattle Times and Newsday lawsuit is the latest battle in a global war between content creators and AI developers. In the United States, the legal framework for addressing these disputes is still being forged. The courts have not yet definitively ruled on whether training AI on copyrighted material constitutes fair use, and the outcome of the New York Times case could set a binding precedent. In Europe, regulators have taken a more aggressive approach, with the EU’s AI Act requiring companies to disclose the copyrighted works used in training and to obtain explicit permission from rights holders in certain circumstances.

The publishing industry is pursuing multiple strategies simultaneously. Some publishers, including The Associated Press and Axel Springer, have chosen to negotiate licensing agreements with OpenAI, receiving compensation for the use of their content. Others, like The New York Times and now The Seattle Times and Newsday, have opted for litigation, seeking to establish a legal principle that AI companies cannot simply take what they want from the internet without paying for it. This divide within the industry reflects broader uncertainty about the best path forward. Licensing deals offer immediate revenue but may legitimize a system that many publishers believe is fundamentally unfair. Lawsuits offer the possibility of a more favorable legal framework but are expensive, time-consuming, and uncertain in outcome.

The technology companies, for their part, have argued that the public interest in AI development justifies the use of publicly available text. They contend that AI systems have the potential to revolutionize education, healthcare, and scientific research, and that overly restrictive copyright laws could stifle innovation. They also point to the technical challenges of licensing every piece of content used in training, given the enormous scale of the datasets involved. OpenAI has offered to allow publishers to opt out of its web crawler, but critics argue that this places the burden on publishers to protect their own content rather than requiring AI companies to obtain permission in advance.

Market Implications: What the Lawsuit Means for the Future of AI and News

The outcome of the Seattle Times and Newsday lawsuit, along with the broader wave of litigation, will have profound implications for both the AI industry and the news business. If the courts rule in favor of the publishers, AI companies may be forced to negotiate licensing agreements with every news organization whose content they use, dramatically increasing the cost of training and potentially limiting the scope of their models. This could reshape the competitive dynamics of the AI industry, favoring companies that have already secured licensing deals and disadvantaging newer entrants that lack the resources to negotiate with thousands of publishers.

For the news industry, a favorable ruling could provide a new revenue stream at a time when traditional advertising and subscription models are under severe pressure. The ability to license content to AI companies could offset declining print revenue and support the continuation of investigative journalism and local reporting. However, there is also a risk that the litigation could accelerate the consolidation of the news industry, with larger publishers benefiting from licensing deals while smaller outlets struggle to enforce their rights.

The strategic significance of the lawsuit extends beyond the immediate parties. The Seattle Times and Newsday are not the largest publishers in the country, but they are representative of a broad swath of regional and local newspapers that form the backbone of American journalism. If these publishers can successfully defend their intellectual property rights, it could embolden hundreds of other local news organizations to pursue similar claims. Conversely, if the courts reject their arguments, it could deal a devastating blow to the economic prospects of local journalism, accelerating the already alarming trend of newspaper closures and news deserts.

To appreciate the legal arguments in the lawsuit, it is helpful to understand the specific mechanics of how large language models are trained. The process begins with data collection, in which automated crawlers scan the internet and download text from websites, including news articles. This text is then processed to remove formatting, extract plain text, and sometimes to filter out low-quality or duplicate content. The resulting dataset, which can be terabytes in size, is used to train the model through a process known as unsupervised learning, in which the model learns to predict the next word in a sequence based on the patterns it has observed.

The critical legal question is whether the intermediate copies made during this process constitute copyright infringement. When a web crawler downloads a news article, it creates a copy of that article on the company’s servers. That copy is then used to train the model, and the model itself can be thought of as a compressed representation of the statistical patterns in the training data. The publishers argue that each of these steps — the downloading, the storage, and the training — involves the unauthorized reproduction of their copyrighted works. The defendants argue that the copies are temporary and incidental to the purpose of creating a non-infringing AI model, and that the model itself does not contain copies of the original works.

The technical reality is more nuanced. While large language models do not store verbatim copies of their training data in the way a database does, they can and do reproduce fragments of that data, particularly for text that appears frequently in the training set. The New York Times demonstrated this by prompting ChatGPT to generate passages that closely matched its articles. The Seattle Times and Newsday lawsuit makes similar allegations, arguing that the models are capable of generating outputs that infringe on their copyrights. The ability of AI models to memorize and reproduce training data is an active area of research, and the extent of this memorization is a key factual question in the litigation.

What Are the Potential Consequences for OpenAI and Microsoft if They Lose the Lawsuit?

If the court rules against OpenAI and Microsoft, the consequences could be severe. The plaintiffs have requested damages for past infringement, which could amount to hundreds of millions of dollars given the number of articles involved. More significantly, they have asked the court to order the destruction of any AI models that were trained on their copyrighted content. Such an order would be extraordinarily disruptive, potentially requiring OpenAI and Microsoft to retrain their flagship models from scratch using only licensed or public domain data. The practical challenges of identifying and isolating the specific data used in training, and of proving that a retrained model does not contain any residual influence from the infringing data, would be immense.

The lawsuit also seeks injunctive relief to prevent future use of the plaintiffs’ content without permission. This could force OpenAI and Microsoft to implement new systems for tracking the provenance of their training data and for obtaining licenses from copyright holders. The cost of compliance would be substantial, and the operational complexity could slow the pace of AI development. For OpenAI, which is already facing a barrage of legal challenges from authors, artists, and publishers, an adverse ruling in this case could compound the company’s legal and financial risks, potentially affecting its valuation and its ability to raise capital.

The Future of Journalism in an AI-Dominated Information Landscape

The lawsuit between The Seattle Times, Newsday, and the AI giants is, at its core, a struggle over the future of information. The plaintiffs are not simply seeking compensation for past harms; they are trying to establish a legal and economic framework that will allow professional journalism to survive in an era when AI can generate text that is increasingly indistinguishable from human writing. The outcome of this case, and the broader legal movement it represents, will determine whether news organizations can continue to invest in the costly and time-consuming work of original reporting, or whether they will be reduced to content farms whose output is fed into AI systems that compete with them for audience attention.

The existential threat that the lawsuit describes is real. The economics of journalism have been deteriorating for two decades, as advertising revenue shifted to digital platforms and readers became accustomed to free access to news. The rise of generative AI threatens to accelerate this decline by offering users a substitute for the original source. If a reader can ask ChatGPT to summarize the news of the day, they have less reason to visit a newspaper’s website or subscribe to its newsletter. The result is a downward spiral in which publishers lose revenue, cut staff, produce less reporting, and become less relevant to the AI systems that once depended on them.

The Seattle Times and Newsday are betting that the courts will recognize the gravity of this threat and will intervene to protect the intellectual property that sustains professional journalism. The metaphor of the snake eating its own tail is a powerful warning about the consequences of inaction. If the AI industry continues to consume journalism without contributing to its production, the quality of both journalism and AI will decline. The question is whether the legal system can act quickly enough to break the cycle before it is too late. The answer will depend on the judges, the juries, and the precedents that emerge from the courtroom battles now unfolding across the country.

Share This Article