Google Defends AI Training as Fair Use in New Policy Paper

In a new policy paper, Google formally argues that training AI on publicly available web content qualifies as fair use under U.S. copyright law.

By Central
Google’s paper frames AI training as a transformative, non-expressive use, drawing a direct analogy to an art student finding inspiration in a museum gallery.
Highlights
  • Google compares AI training to an art student learning from gallery works, calling it a non-expressive use of data.
  • The company recommends publishers use robots.txt and Google-Extended controls to opt out of training data.
  • Google leaves the door open for future paid access arrangements without committing to specific terms or timelines.

Google has formally staked its position on one of the most consequential questions facing the AI industry: whether training large language models on publicly available web content constitutes copyright infringement. In a policy paper published June 25, the company argues that such training qualifies as fair use under U.S. law, drawing a direct analogy to an art student finding inspiration in a museum gallery. The paper, titled “A Pragmatic Approach to AI Governance in America,” consolidates Google’s previously stated views and arrives amid intensifying pressure from publishers and regulators who are demanding clearer boundaries around AI’s access to copyrighted material. For search professionals, publishers, and content creators who have been watching the AI training debate closely, the document offers the clearest articulation yet of Google’s official stance on how it believes AI systems should interact with the open web.

The paper lands at a moment when the relationship between AI companies and content publishers is under unprecedented scrutiny. Since the launch of AI Overviews, the question of how models are trained has moved from a technical backend concern to a central public policy issue. Google’s response is to frame its training methods as a transformative, non-expressive use of data that should remain protected under fair use doctrine, while pointing to existing opt-out mechanisms and copyright law as sufficient safeguards for publishers who do not want their content used. The company also acknowledges, however, that specific arrangements for paid access to specialized content may be appropriate in certain circumstances, leaving the door open to future commercial agreements without committing to any particular program or timeline.

Google’s central argument is that training AI models on publicly available web data is a transformative use that does not infringe copyright. The company draws a direct comparison to an art student who walks through a gallery and takes inspiration from the works on display, suggesting that the act of learning from publicly accessible content is fundamentally different from reproducing that content. In legal terms, Google describes the training process as a non-expressive use of data, meaning the model does not communicate the original works to users but rather learns patterns and relationships from them.

What is Google’s position on AI training and fair use? Google argues that training AI models on publicly available web data qualifies as fair use under U.S. copyright law because it is a transformative, non-expressive use of the data. The company maintains that this approach should be protected as fair use and that similar protections should be extended internationally through text-and-data-mining exceptions in copyright law.

For publishers who do not want their content used for training, Google recommends using machine-readable controls such as Google-Extended in their robots.txt files. This places the responsibility on content creators to proactively signal their preferences rather than requiring AI companies to seek permission before using the content. When AI outputs do produce content that closely resembles existing copyrighted works, Google argues that the existing notice-and-takedown framework is sufficient, rather than implementing proactive filtering systems that would judge whether an output is too similar to the source material.

The Opt-Out Model and Its Growing Challenges

The opt-out approach that Google advocates has become a flashpoint in the broader debate over AI and copyright. Publishers and industry groups are increasingly pushing back against the idea that content owners should bear the burden of requesting exclusion from AI training datasets. Digital Content Next, a trade association representing major publishers, recently sent a cease and desist letter to the Common Crawl Foundation, arguing that copyright law is not an opt-out regime. The letter emphasized that scrapers should seek permission before using content, rather than requiring publishers to affirmatively request to be excluded. This perspective directly challenges the foundational assumption of Google’s policy paper, which treats opt-out controls as the primary solution for publisher concerns.

Google has already begun testing an opt-out toggle for AI search features, following regulatory pressure from the UK’s Competition and Markets Authority. The CMA introduced a new conduct requirement this month that gives websites the option to opt out of AI search features and requires Google to attribute publisher content. The regulator explicitly stated that this measure is intended to help boost publishers’ bargaining power. However, the reports available to publishers to help them make informed decisions about opting out have not yet included click data, limiting their ability to assess the true impact of AI search features on their traffic and revenue.

Regulatory Momentum and Publisher Pressure

The policy paper arrives at a time when regulatory momentum is building on multiple fronts. In the UK, the CMA’s new requirements represent a significant step toward giving publishers more control over how their content appears in AI-powered search features. The attribution requirement, in particular, addresses a key publisher concern that AI systems often use content without providing clear credit or linking back to the original source. Google has framed its own approach as compatible with these requirements, pointing to the opt-out controls it is developing and the attribution mechanisms already built into its search products.

US publishers are making their stance even clearer and more aggressive. Digital Content Next’s cease and desist letter to Common Crawl reflects a growing frustration with the assumption that publicly available content is freely available for AI training. The letter’s argument that copyright law is not an opt-out regime represents a fundamental challenge to the approach that Google and other AI companies have been advocating. If this view gains traction with regulators or in the courts, it could fundamentally reshape how AI companies access and use web content for training purposes.

Google’s paper acknowledges these tensions but does not offer significant concessions. The company continues to advocate for keeping its current approach largely unchanged, emphasizing that existing legal frameworks and opt-out mechanisms are sufficient to address publisher concerns. For regulators considering new rules, the paper makes clear that Google believes the current balance between AI innovation and copyright protection is already appropriate and that significant new restrictions could harm the development of AI technologies.

What the Paper Reveals About Google’s Strategy

Beyond its specific arguments about fair use and opt-out controls, the policy paper reveals important aspects of Google’s broader strategy for navigating the AI copyright landscape. The company is positioning itself as a defender of the open web’s accessibility for AI training purposes, arguing that overly restrictive copyright rules could undermine the development of beneficial AI systems. At the same time, Google is signaling a willingness to explore commercial arrangements for certain types of content, particularly specialized or non-public data that could help keep AI responses up-to-date and accurate.

The mention of grounding partnerships and content deals is significant, though the paper provides no specifics about which programs, terms, or timelines Google has in mind. This ambiguity leaves room for Google to develop customized arrangements with individual publishers or content providers, potentially offering payment for access to high-value content while maintaining the general principle that most publicly available data is free for training purposes. The paper’s language suggests that Google sees value-exchange relationships as appropriate for certain contexts but not as a universal requirement for all content used in AI training.

For publishers trying to understand how to manage AI access to their content, the paper offers a clear picture of Google’s starting position in negotiations. The company is advocating for the maximum possible freedom to use publicly available content for training, with opt-out controls as the primary mechanism for publishers who want to restrict access. Any movement beyond this position, whether toward permission-first scraping, compensation requirements, or more detailed attribution standards, will likely require additional regulatory pressure or court rulings that shift the legal landscape.

Implications for Publishers and the Search Ecosystem

For publishers and search professionals, the policy paper provides both clarity and a basis for planning. Google’s position is now formally documented and publicly available, removing any ambiguity about where the company stands on the core questions of copyright and AI training. Publishers who are concerned about their content being used for training can implement Google-Extended in their robots.txt files as a signal of their preferences, though the effectiveness of this approach depends on whether other AI companies also respect such signals and whether the opt-out model holds up under legal and regulatory scrutiny.

The broader implications extend beyond individual publishers’ decisions about whether to opt out. If the opt-out model becomes the standard approach for AI training, it could create a two-tier system in which publishers who allow their content to be used for training gain certain advantages, such as potentially better representation in AI responses or access to partnership programs, while those who opt out maintain stricter control over their intellectual property. This dynamic could create new strategic considerations for publishers trying to balance the benefits of AI visibility against the risks of having their content used without direct compensation.

The regulatory landscape will continue to evolve, and Google’s paper is as much a lobbying document as it is a policy statement. By publishing its position now, Google is seeking to influence the terms of the debate before regulators and courts make decisions that could constrain the company’s approach. The paper’s emphasis on the international dimension, particularly its call for text-and-data-mining exceptions to be extended globally, reflects an understanding that the outcome of this debate in the US and UK will set precedents that could ripple across other jurisdictions.

The tension between Google’s position and the demands of publishers and regulators is unlikely to resolve quickly. The paper offers controls and signals a willingness to negotiate individual arrangements, but it does not concede the fundamental principle that training on publicly available data should require permission or compensation. As policymakers in the US, UK, and other jurisdictions consider new rules for AI and copyright, Google’s paper will serve as a reference point for what the company considers reasonable and defensible, while publishers and their advocates will continue to push for a framework that starts with permission rather than opt-out.

The details of Google’s partnership programs, content deals, and any associated payment structures remain unspecified, which means the practical consequences for publishers will depend on how these elements develop in practice rather than on the policy positions outlined in the paper. The value-exchange language in the document is notable, but it will be meaningful only if it translates into concrete programs with clear terms and tangible benefits for content creators. For now, the paper stands as a carefully crafted articulation of Google’s ideal regulatory environment, one in which AI training continues under fair use protections, opt-out controls suffice for concerned publishers, and commercial arrangements emerge organically for specialized cases where both parties see mutual benefit.

Share This Article