Google has begun rolling out its June 2026 spam update, the second such enforcement action of the year, and with it comes a freshly sharpened policy: attempts to manipulate generative AI responses in Search are now explicitly classified as spam. The language in Google’s documented spam policies has been expanded to cover this territory, and the update is actively enforcing that boundary. But almost as soon as the policy landed, a new preprint from Cornell Tech, highlighted by 404 Media, laid bare why enforcement will be far harder than the policy’s clean wording suggests. The community pages that AI research agents depend on to produce answers can carry third-party comments, and a single planted comment can inject a recommendation that the original author never wrote. What Google labels spam therefore travels through the very retrieval pipelines these agents rely on, and the obvious defenses all carry significant drawbacks.
What the June 2026 Spam Update Actually Enforces
The June spam update enforces Google’s documented spam policies, and the key change is that those policies now treat the manipulation of generative AI responses as a violation. This means that tactics designed to influence what appears in AI Overviews, AI Mode, or any Google-owned AI search surface are subject to the same penalties as traditional link spam, cloaking, and scraped content. Google’s SpamBrain system and manual review teams are the primary enforcement mechanisms, though the company has not detailed how specifically it will detect generative-AI manipulation at scale.
For anyone trying to push a brand into AI-generated answers, the line between legitimate optimization and prohibited spam is being redrawn in real time. The policy names the behavior, but as the Cornell research demonstrates, the infrastructure of modern AI search makes it nearly impossible for the site involved to know it has been used as a vector.
Why the Policy Is Harder to Enforce Than It Looks
Google’s spam rules treat attempts to “manipulate generative AI responses” as a violation, and the June update enforces that policy. But a Cornell Tech preprint picked up by 404 Media gets at why enforcement is so difficult. The community pages that AI research agents lean on can also carry third-party comments, and a comment can plant a recommendation that the author never wrote.
The paper, titled “Deep-Research Agents Can Be Poisoned via User-Generated Content,” has not yet been peer-reviewed. It probes a specific weak spot in how AI research tools collect their sources. These tools answer a question by firing off a batch of related sub-queries, grabbing the pages that keep coming up across them, and assembling a report with citations. Analysis revealed that the same community pages surface repeatedly in those sub-queries. Inside a single topic cluster, one user-generated page turned up in as many as 48 percent of queries, and user-generated platforms made up 17 to 23 percent of every URL retrieved. Alter one of those recurring pages, and the change can ripple into the reports for an entire topic.
The authors found that roughly 13 words of planted text on a recurring page were enough to insert an attacker’s chosen entity into the finished report in 38 to 51 percent of sessions that retrieved the page. Scatter the same text across a handful of pages, and the figure climbed to 42 to 62 percent. Even buried inside a full page, where it made up under 4 percent of what the agent read, the planted text still surfaced in 30 to 53 percent of sessions. Three open-source research agents took the tests—STORM, Co-STORM, and OmniThink—all run in a simulation so that nothing on the live web was touched.
The Stakes for Brands and Publishers
SE Ranking’s tracking of AI Mode found Google increasingly pointing to its own properties, with self-citations rising to roughly a fifth of all AI Mode citations in its latest report. With more citations pointing to Google and fewer to external websites, the pull to manufacture one rises accordingly. A gray market has already begun to form, and the Cornell authors point out that marketers are busy testing ways to nudge AI-generated answers.
Businesses, meanwhile, do not have the data they need to see what is happening. No dashboard tells a site whether it landed in an AI answer, got cited in a generated report, or was passed over. The result is a violation Google can name but the site involved often cannot see.
What Is AI Answer Manipulation and How Does It Work?
AI answer manipulation refers to the practice of planting text across user-generated platforms—such as Reddit, forums, and review sites—with the specific intent of influencing what appears in AI-generated search summaries or research reports. The Cornell research demonstrates that roughly 13 words of strategically placed text on a recurring community page can successfully inject a chosen entity into an AI report in up to 51 percent of retrieval sessions. The planted text reads like real advice, sits on the same pages the tools were always going to read, and telling it apart from a normal post is the core challenge for any enforcement system.
Where the Research Shows Defenses Fall Short
The research team looked for a defense against planted text but did not find one that worked without degrading the user experience. They tried cutting user-generated sources out entirely, screening them with a language model before use, and combing the finished report for claims that did not hold up. None of the three approaches stopped the attack without making the results worse for the user. Drop the user-generated sources, and you lose the community detail that makes AI search tools worth using.
The tools most people use sit outside that test. ChatGPT Deep Research and Gemini Deep Research run retrieval the researchers could not poison without crossing an ethical line, so they only measured citation habits. Gemini leaned on user-generated content 12.1 percent of the time, which the authors call a hint of exposure, not a tested result. OpenAI’s tool reached for it far less.
Why This Update Matters for Search Professionals
The moves that can help lift a brand into AI answers are disturbingly similar to the manipulation tactics Google now calls spam—namely, planting mentions across the sites these tools read. No one outside Google knows where the line falls between earning a mention and engineering one.
For ecommerce and local brands, the danger comes from the other direction. The test cases in the Cornell research were the ordinary things people ask: which service to call, which product to buy, where to eat. A rival or a scammer can slip an unfamiliar name into those answers, right next to the legitimate options, and the brand being edged out would never know it.
For news publishers and larger brands, the worry is trust in the answer their name lands in. A citation from an AI tool is widely seen as a win, but a citation only reflects what the tool pulled, not whether that page was accurate, and the answer can be steered by content the brand never wrote. There is no tidy fix to all this. AI visibility has become a surface you actively monitor, not just a channel you passively optimize for.
The Open Problem No Single Platform Can Solve Alone
The authors called user-generated manipulation an open problem that no single platform can fix on its own. Reddit has flagged its long-running fight against coordinated manipulation, and Google has bolted context labels onto some Reddit-sourced material in AI Overviews. Neither approach touches the retrieval concentration the paper points to—the fact that a single user-generated page can appear in nearly half of all sub-queries for a given topic cluster.
Google has not indicated how it intends to enforce generative-AI manipulation at scale, whether through a dedicated update, its SpamBrain system, or manual reviews. For now, the policy calls the behavior out of bounds, and vetting AI responses still rests with whoever is reading them. The June 2026 spam update draws the line, but the research makes clear that holding it will require far more than a policy change.