US DOJ backs fair use for AI training in copyright case

The US Department of Justice intervenes in a landmark AI copyright case, arguing that training models on copyrighted material is fair use.

By Central
The DOJ's filing distinguishes between training and output, rejecting the New York Times' claim of infringement.
Highlights
  • The DOJ argues that copying during AI training is transformative and does not constitute infringement.
  • The department compares AI training to a student reading books or copying texts to learn.
  • The outcome will shape how copyright law applies to AI and the future of knowledge creation.

The United States Department of Justice has intervened in one of the most consequential legal battles for the future of artificial intelligence, arguing that training AI models on copyrighted material constitutes fair use. The filing, submitted in the consolidated lawsuit against OpenAI and Microsoft led by the New York Times, rejects the core premise that ingesting copyrighted works to teach a machine constitutes infringement.

The case, lodged in a Manhattan federal court in late 2023, alleges that OpenAI and Microsoft used millions of New York Times articles without permission to train large language models like GPT-4. The newspaper claims these models now compete directly with its own journalism as a source of information, seeking billions of dollars in damages and demanding the destruction of the models allegedly built on its content. The dispute has escalated into a bellwether for how American courts will reconcile traditional copyright law with the mechanics of modern AI training, making the DOJ’s intervention a pivotal moment.

The DOJ Draws a Line Between Training and Output

The core of the Justice Department’s argument rests on a distinction between the act of training and the act of generating. During the training phase, entire works are copied, but those copies are never made publicly available. The outputs of the model, the DOJ contends, “often if not always lack substantial similarity” to the original training material. This distinction is legally critical under the fair use doctrine, which considers the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the potential market.

The DOJ argues that a blanket theory of market harm, which conflates the creation of a training dataset with the public release of a competing product, is legally unsound. It is not enough, the filing suggests, to merely assert that a model could theoretically generate content that harms the original market. The harm must be tied directly to the specific use of the copied material, and in the case of training, the department sees that link as broken.

What Is the Legal Basis for the DOJ’s Fair Use Argument in AI Training?

The DOJ’s argument hinges on the transformative nature of the copying during model training. The department compares the process to a teenager copying Hemingway’s stories by hand to learn how his sentences worked. The filing invokes an analogy to Joan Didion, who famously used this method. Under the logic the DOJ is contesting, Didion could have faced copyright liability whenever she later published her own writing, because her learning process and subsequent commercial work would be treated as a single, infringing use. The department contends that requiring payment every time a person draws on a learned work to create something new would be fundamentally at odds with the purpose of copyright, which is to promote creativity.

The filing directly cites the reasoning from the Kadrey ruling, which dealt with training models on books. The DOJ also leans on an earlier ruling that argued it would be unthinkable to require people to pay a license fee whenever they later draw on a book to write something novel. The department frames the AI training process as a mechanical, non-expressive act of learning, distinct from the public distribution of substantially similar works.

The DOJ’s position places it in direct conflict with the US Copyright Office. In a recent report, the Copyright Office explicitly rejected blanket fair use for AI training, arguing that the technology operates at a speed and scale far beyond human capabilities, generating content that can directly compete with original works in existing markets. The Copyright Office found that commercial applications, especially those designed to replace the very sources they learn from, exceed what fair use allows under the law.

The DOJ goes on the offensive against that report, arguing that the assessment from former Copyright Register Shira Perlmutter carries no binding legal authority. The department claims the report ignored settled case law regarding case-by-case analysis and misidentified the types of market harm that are legally relevant under the statute. The filing doubles down on the principle that each use must be evaluated individually, a position that implicitly rejects the Copyright Office’s broader, categorical warning.

The timing of the intervention is politically charged. Perlmutter was fired by the Trump administration shortly after her report was published. New York Democrat Joe Morelle stated that she was let go because she refused to legitimize AI training on copyrighted works, a position reportedly favored by Trump ally Elon Musk. In a footnote within its filing, the DOJ acknowledges that Perlmutter is currently challenging her dismissal, highlighting the political tensions underlying what is ostensibly a legal debate.

How Does Scale Change the Fair Use Analysis for AI?

The DOJ’s arguments, however incisive they may be in isolation, are open to a significant counterpoint: scale. A single author copying a few pages of Hemingway to learn sentence structure is a fundamentally different act from a multibillion-dollar corporation copying terabytes of text to build a mass-market product. The US Copyright Office made this point explicitly, noting that while the copying done by a human for personal study might be fair use, the same wholesale copying by a company for commercial AI training pushes beyond traditional limits.

The DOJ’s analogy to Joan Didion, while rhetorically powerful, glosses over the mechanical reality of AI training. Didion copied Hemingway deliberately, thoughtfully, and at a human scale. An AI training run copies without comprehension, on a scale that dwarfs any conceivable human effort. The output, while not a direct copy, is a statistical remix of millions of texts, capable of reproducing facts, styles, and structures from specific sources on demand. The question the courts must answer is whether this quantitative difference amounts to a qualitative difference in legal terms.

The DOJ seems to acknowledge this tension by arguing that even if some uses of a model’s output are infringing, the training process itself should not be presumptively illegal. The department is trying to protect the foundational act of machine learning, leaving the door open for lawsuits against specific, harmful outputs. This is a nuanced position: protect the means of creation, but still hold creators accountable for what they produce.

Market Harm and the Future of News

For the New York Times and other publishers, the central issue is market substitution. They argue that AI models, particularly when used in search or as chatbot replacements, directly siphon off readers who would otherwise visit their websites and generate advertising revenue or subscription fees. The newspaper’s lawsuit is not just about past copying; it is about the future of its business model. If a user can ask a chatbot a question and receive a synthesized answer built from dozens of paywalled articles, the publisher loses that traffic and that revenue.

The DOJ counters that this theory of market harm is too broad. The department argues that the relevant market harm under copyright law must be tied to the specific copying that occurred during training, not to the general competition that the resulting product provides. An encyclopedia does not infringe on the copyright of every book it summarizes, the logic goes, even if it replaces the need to buy those books. The DOJ is pushing the court to view the training data as raw material and the model as the new, transformative work.

The filing also points out an irony that touches on the very journalists bringing the suit. The DOJ notes that writers at the New York Times reportedly use LLMs to draft and edit articles. While the link provided by the department largely supports a more cautious conclusion about this practice, the point stands: the industry is already intertwined with the technology it is suing to limit. This creates a complex legal landscape where the plaintiffs are both competitors and potential users of the very tools they seek to restrict.

Broader Implications for the AI Industry

The DOJ’s intervention provides a significant, though not decisive, boost to AI companies like OpenAI, Microsoft, and Anthropic, which are facing a wave of similar lawsuits from authors, visual artists, music labels, and news organizations. A ruling in favor of the DOJ’s position would create a powerful precedent, effectively greenlighting the current training practices of the industry. It would mean that the primary legal risk for AI companies lies not in how they build their models, but in how they deploy them.

Conversely, a ruling against the DOJ’s position could have a chilling effect. It could force companies to license every piece of data they use for training, a logistical and financial nightmare that would likely favor large, established players who can afford massive licensing deals. It could also lead to the destruction of existing models, as the New York Times has demanded, setting a precedent that the cost of innovation is the risk of total loss if the legal ground shifts.

The DOJ filing is clear that its position is not an endorsement of unlimited copying. The department argues for a legal framework that allows training to proceed while preserving the right to sue over specific, infringing outputs. This distinction is critical for the development of AI safety and alignment research. If the training process is legal, researchers can continue to build and study models as foundational tools, but they must be extremely careful about how those tools are packaged and sold to the public.

A Question of Creativity and Control

At its heart, this legal battle is a fight over the definition of creation. The DOJ argues that creation is a human activity, and that AI models are merely tools that humans can use to create. Training those tools on existing works is akin to a student reading a library full of books. The publishers argue that creation requires some form of licensing or permission, and that unlicensed, massive-scale copying for profit is theft, not learning.

The DOJ explicitly states that “human beings create original works using LLMs” and that imposing liability for training data would stifle the very creativity copyright law is intended to protect. This is a forward-looking, almost philosophical claim. It asserts that the future of creativity depends on access to the past, and that locking up data behind licensing walls would impoverish culture rather than enrich it. The publishers, conversely, argue that without control over their own property, they cannot invest in high-quality journalism, and that a parasitic ecosystem of AI remixes will ultimately destroy the sources of original information.

The court will have to decide which vision aligns with the text of the Copyright Act and the precedents of decades of fair use jurisprudence. The DOJ has made its position clear: the act of teaching a machine is not the same as the act of publishing a copy, and the law must adapt to protect the process of learning, even when that learning happens at silicon scale. The ultimate outcome will shape not just the business models of AI companies and publishers, but the very nature of how knowledge is aggregated, synthesized, and created in the digital age.

Share This Article