For more than two years, opposing armies have been massing and digging in as they prepare for a battle royale to decide the future of their respective domains. No, the Iron Throne does not hang in the balance, but the future of generative artificial intelligence, and its potential to better the lives of literally everyone everywhere, does. Those are the stakes in a massive consolidated copyright legal case against OpenAI brought by dozens of authors, including George R.R. Martin, and major news organizations. After years of preliminary jousting by lawyers and experts, the central legal issue will come to a head in the fall and is expected to be argued early next year before a federal judge in New York City. The outcome will determine whether fair use for AI training survives, and with it the accessibility, affordability, and diversity of the most transformative technology since the internet.
What Is Fair Use and Why Does It Matter for AI Training?
Fair use is a cornerstone of American copyright law, established long before the Civil War and codified into the Copyright Act by Congress 50 years ago this fall. It expressly allows copyrighted works to be used without the prior permission of their owners to create something fundamentally new and different, provided the new work does not substitute for the original itself. For generative AI, this doctrine is the legal foundation upon which developers have trained large language models like ChatGPT on vast swaths of text, images, and data drawn largely from the open internet. Without fair use, every piece of copyrighted material used in training would require individual licensing agreements, a logistical and financial impossibility at the scale required for modern AI systems.
The irony is palpable: the very authors and news outlets now suing to block AI training regularly rely on fair use of copyrighted material to write their own books and articles. Journalists quote from competing publications, historians excerpt primary sources, and novelists reference earlier works all without seeking permission. Fair use is not a loophole; it is a deliberate, constitutional mechanism designed to ensure that copyright law promotes progress rather than stifles it. Yet plaintiffs in the consolidated case are asking courts for orders that would force developers to pull existing generative AI models off the market entirely until they can be retrained using only licensed material.
Developers like OpenAI push back with a different reading of the law. They argue that their use of copyrighted material is a textbook case of fair use under decades of precedent, because the material on which gen AI models are trained is not copied, stored, or resold. Instead, it is transformed into a new kind of tool, not a substitute for any specific book or article. They also cite a clear line of Supreme Court and appellate rulings that the maker of a multi-purpose product is not usually on the hook when someone else uses that product to violate copyright law. The legal question is narrow, but its implications are vast.
How the Fair Use Battle Could Reshape the AI Industry
For most of the general public, this may sound like an obscure legal fight between powerful companies over money, but it is much bigger than that. If courts dramatically narrow fair use for AI training, the consequences would reach every classroom, laboratory, startup, library, marketplace, and workplace that increasingly relies on generative artificial intelligence. The ruling would not just affect OpenAI and a handful of tech giants; it would fundamentally alter who can participate in the AI economy and at what cost. Four distinct and damaging outcomes are likely if the plaintiffs prevail.
1. Innovation Would Become Harder, and the Next Generation of AI Companies Might Never Get Started
Today leading AI models require enormous amounts of information during training. The alternative to using fair use material would be to buy or lease that material. It sounds simple enough, but the volume of material required is astronomically large and would be astronomically expensive. The cost and mechanics of simply identifying and negotiating for the huge amount of data necessary would itself be prohibitive for all but the biggest existing developers. For startups and academic labs with limited budgets, this is a non-starter. The result would be a market dominated by a handful of incumbent corporations that can afford to license data at scale, choking off the pipeline of new entrants, new ideas, and new approaches that have driven AI progress over the past decade.
This is not a hypothetical. The history of technology is littered with innovations that emerged from small teams working with limited resources, from the personal computer to the web browser to open-source machine learning frameworks. A licensing-only regime for training data would effectively lock the door on the next generation of AI builders before they even begin.
2. Generative AI Would Become More Expensive and Less Accessible
Requiring gen AI developers to use only licensed training material would invariably increase the cost of developing new models. These costs would ultimately flow to users, affecting not only commercial enterprises but also schools, libraries, other public institutions, governments of every size, and individual consumers. As developers unable to absorb all of the substantially heightened training costs struggle to compete, all users should expect higher subscription costs and likely lower usage limits. The largest corporations would still be able to afford premium AI systems. Smaller businesses, nonprofits, and public institutions would have fewer choices and fewer capabilities.
The accessibility gap would be felt most acutely in education. School districts already operating on tight budgets would face a choice between paying for AI tools or investing in teachers, textbooks, and infrastructure. Public libraries, which have historically served as equalizers of access to information, would find themselves priced out of the most powerful research assistants ever built. Governments at the municipal and state level, which could use AI to improve public services, would be forced to ration usage or forgo the technology altogether. The digital divide would widen into a chasm.
3. Scientific Research, Technical Innovation, and Knowledge Discovery Would Slow
If access to training data became substantially more restricted, specialized research models often built from general-purpose gen AI models would likely become more difficult and more expensive to build. The pace of scientific progress would slow, not because researchers lacked ideas, but because they lacked affordable tools. Consider the fields of drug discovery, climate modeling, and materials science, where AI is already accelerating breakthroughs by analyzing vast literatures and generating novel hypotheses. These applications depend on models trained on broad, diverse datasets that include copyrighted scientific papers, textbooks, and technical manuals. Restricting training data would force researchers to either pay exorbitant licensing fees or work with smaller, less representative datasets, both of which would degrade model performance and slow discovery.
How Would the Public’s Access to Knowledge Be Limited and Fragmented?
The public’s access to knowledge would become limited and more fragmented. Perhaps the least appreciated consequence of eliminating fair use for AI training is how it would restrict and segment the knowledge available to everyone through future AI systems. Today’s best models owe their benefits directly to training on extraordinarily broad collections of books, newspapers, academic writing, historical documents, and countless other sources. If every copyright owner could decide whether its works were used in gen AI training, future models would inevitably become patchworks with a fraction of their former power and utility. In that world, some models might include particular newspapers but not others, contemporary books but not historical archives, and major commercial publications but not local and independent reporting. For ordinary users, this would mean less complete answers about current events, history, literature, and specialized subjects; greater dependence on paywalled material and multiple subscriptions; potentially varied answers to the same inquiry depending on the user’s employer, school, or subscription level; reduced coverage of local, niche, minority-language, and out-of-print material; and a widening gap between well-funded and underfunded educational institutions.
The fragmentation of knowledge is not merely an inconvenience. It is a threat to democratic discourse and informed citizenship. When different populations have access to different quality of information through their AI tools, shared understanding becomes harder to achieve. When answers vary based on who you are or where you study, trust in the technology erodes. The promise of generative AI has always been its ability to democratize expertise, to put a world of knowledge at everyone’s fingertips. A licensing-only regime would turn that promise into a tiered subscription service, with the best answers reserved for those who can pay the most.
Here There Be Dragons: The Real Danger to Copyright
Ancient maps of the globe advised navigators to avoid uncharted waters by marking them with the often appropriately illustrated warning: “Here there be dragons.” The plaintiffs in these cases ask courts to believe that gen AI has taken us into exactly that kind of unexplored and dangerous legal territory. They portray AI training as a novel threat to copyright that requires new restrictions and new remedies. But the truth is that we have a clear map that has provided copyright holders with strong rights and given all of us maximal benefit from incredible new technologies. That map is the Constitution, which says unambiguously that the purpose of copyright law is “to promote the progress of science and useful arts.”
Robust fair use has been and remains the principal way U.S. copyright law strikes the balance between compensation and innovation. That balance must continue. The real danger to copyright is not from generative AI. It is posed by fire-breathing plaintiffs intent on impeding a new technology’s development by incinerating the foundations of fair use. A ruling that dramatically narrows fair use for AI training would not protect authors; it would protect incumbent market positions. It would not preserve the value of creative work; it would ration access to knowledge. It would not uphold the Constitution’s mandate to promote progress; it would undermine it.
The coming courtroom battle in New York City is more than a legal dispute. It is a referendum on whether the United States will remain the world’s leader in artificial intelligence and whether the benefits of that technology will be broadly shared or concentrated among the wealthy and powerful. The stakes could hardly be higher. The fair use doctrine has served the nation well for more than a century and a half, enabling everything from the printing press to the internet. It must now be allowed to do the same for generative AI.