Hidden Prompts Trick AI Into False Email Summaries

New research reveals how attackers exploit AI summarization tools to generate false email summaries with 100% success.

By Central
Forcepoint X-Labs demonstrates indirect prompt injection in email HTML, achieving a 100% success rate across ten trials.
Highlights
  • Attackers embed invisible instructions in email HTML to force AI into generating false summaries.
  • Forcepoint X-Labs achieved a 100% success rate in ten trials of prompt injection attacks.
  • OWASP has ranked prompt injection as the top risk for LLM applications since 2023.

Hidden Prompts Trick AI Into False Email Summaries

A new wave of cyber attacks is exploiting a fundamental flaw in how large language models process information, and the consequences are already being measured in manipulated invoices and altered business communications. Researchers at Forcepoint X-Labs have demonstrated that attackers can embed invisible instructions within the HTML of an email to force AI-powered summarization tools into generating false summaries that appear completely legitimate to the recipient. The technique, a variation of indirect prompt injection, succeeded in 100 percent of test trials, raising urgent questions about the security of AI systems now widely deployed to read, summarize, and act on corporate email.

The study marks the latest and most precisely measured demonstration of a vulnerability that security experts have warned about for years. The core problem is that AI systems, including the most advanced large language models, cannot reliably distinguish between the data they are meant to process and the instructions they are meant to follow. When an AI assistant reads an email to produce a summary, it treats the entire content of that email—including any hidden commands—as part of its input. If an attacker can embed a prompt inside that email, the AI may follow the attacker’s instructions rather than the user’s intent.

This is not a theoretical concern. The Open Web Application Security Project (OWASP) has consistently ranked prompt injection as the number one risk for LLM and generative AI applications in its Top 10 list since 2023. The problem is so pervasive and so difficult to solve that it has become the defining security challenge of the generative AI era. Forcepoint’s new research provides concrete evidence of how easily this vulnerability can be weaponized in a real-world communication channel—corporate email—and what organizations must do to defend against it.

How Attackers Hide Malicious Prompts in Plain Sight

The Forcepoint X-Labs team constructed a proof-of-concept experiment designed to simulate a realistic corporate email environment. They built an isolated lab with synthetic data and an Outlook add-in that forwarded email headers and body text to an LLM-powered summarization service. The model chosen for the test was Claude Haiku 4.5, a widely used AI system capable of generating concise and accurate summaries under normal conditions.

The researchers deliberately created a simple email-to-LLM pipeline. There were no guardrails, no safeguards, and no mechanisms that would allow the AI to distinguish between the email’s legitimate content and any embedded instructions. This is not an unrealistic scenario. Many organizations have deployed AI summarization tools with minimal security considerations, often as quick integrations with existing email platforms.

Into this pipeline, the researchers injected a seemingly normal email. The email contained a malicious prompt hidden within the HTML code. The trick was deceptively simple: the researchers used a font size and color combination that rendered the hidden text invisible to the human recipient viewing the email in Outlook. The invisible text remained present in the HTML source that the email summarizer received. The AI, processing the raw HTML, read the hidden prompt as part of the email content and acted on it.

100% Success Rate: The Numbers Behind the Attack

To establish a baseline, the researchers submitted a clean version of the test email to the summarizer ten times. The AI generated accurate summaries in all ten runs. Then they submitted the injected version of the same email, again ten times. The injection succeeded in all ten runs. In every single instance, the AI generated a summary that contained altered information, and it did so without any indication that the summary had been compromised.

The specific manipulations were subtle but significant. The original email mentioned an outstanding invoice amount of €8,750. The injected version produced a summary showing an outstanding amount of €46,200—a difference of more than five times the actual figure. Similarly, the original email referenced a date for a made-up quarterly supplier review. The injected version shifted that date entirely. A recipient relying on the AI summary would have acted on completely false information without any visible warning.

Ben Gibney, the Forcepoint researcher who led the study, emphasized the controlled nature of the experiment. “This is a simple test with only one message, using one model, and a single run of ten trials each for the benign and injected emails,” Gibney wrote. “It’s not a full attack scenario with thousands of similar messages targeting many victims.” The implication is clear: the vulnerability exists, it works reliably, and scaling it to a broader attack is a matter of attacker resources, not technical feasibility.

Why AI Cannot Distinguish Data from Instructions

What is prompt injection? It is a class of attack in which an attacker crafts input that causes an AI model to behave in unintended ways. The input appears to the model as legitimate data, but it contains instructions that override or alter the model’s intended behavior. In the case of email summarization, the attacker’s hidden prompt might instruct the AI to ignore certain parts of the email, fabricate numbers, change dates, or even generate requests for payment that never existed in the original message.

The reason this works is architectural. Large language models process text as a sequence of tokens. They have no inherent mechanism to distinguish between tokens that represent user instructions and tokens that represent untrusted data from an external source. When an AI is given the task of summarizing an email, it must read the entire email to understand its content. If that email contains a hidden instruction, the model treats it as part of the same input stream. The model has no way to know that the instruction came from an attacker rather than the user.

This is not a bug that can be patched with a simple update. It is a fundamental limitation of current AI architectures. Researchers have proposed various mitigation strategies, including input sanitization, prompt isolation, and output verification, but none have proven foolproof. The OWASP ranking reflects the industry’s consensus that prompt injection is not a temporary vulnerability but an enduring risk that organizations must learn to manage.

The Greater Danger: Agentic Summarizers That Can Act

Forcepoint’s experiment focused on a relatively simple scenario: an AI that reads an email and generates a summary. The manipulated summaries contained false information, but the AI itself did not take any further action. Gibney pointed out that the security implications become far more severe when the AI summarizer has agentic capabilities—meaning it can perform actions based on what it reads.

“The security implications are bounded to what the summarizer is directed to do,” Gibney told Dark Reading. “But an agentic summarizer given the ability to send emails, schedule meetings and more, would have much greater security implications.”

Consider an AI assistant that not only summarizes emails but also drafts replies, sends calendar invitations, or initiates payment approvals. An attacker who can inject a hidden prompt into an email could potentially instruct the AI to send a phishing reply to a colleague, schedule a fake meeting with an external party, or generate a fraudulent payment request. The AI would carry out these actions with the same apparent legitimacy as any other task, and the human user might never suspect that the original email contained malicious instructions.

The industry is moving rapidly toward agentic AI systems. Major technology companies are building AI assistants that can browse the web, interact with APIs, and execute commands on behalf of users. Every one of these systems must process external content to function. And every one of them is potentially vulnerable to the same class of attack that Forcepoint demonstrated with a simple email summarizer.

Forcepoint’s Contribution: Measurement, Not Discovery

Gibney was careful to frame the research in context. Prompt injection is not a new discovery. Researchers have demonstrated various forms of the attack for several years, and the technique of hiding text in HTML emails is decades old. What Forcepoint’s study adds is rigorous measurement. By pre-registering the facts in both the benign and injected emails—meaning they specified exactly what information should appear in a correct summary and what should appear in a manipulated one—the researchers were able to confirm with statistical certainty that the injection succeeded every time.

“We identified this by pre-registering the facts in both the benign and injected emails, then confirmed the injections held,” Gibney explained. This methodological rigor matters because it moves the conversation from “this attack is possible” to “this attack works with 100 percent reliability under test conditions.” Organizations can no longer treat prompt injection as a theoretical risk. The data shows it is a practical and repeatable threat.

Structural Defenses: Separating Trusted from Untrusted

The most important recommendation from Forcepoint’s research is that organizations must treat incoming content and AI-generated output as potentially untrusted. This is a fundamental shift in mindset for many security teams, who are accustomed to trusting internal email traffic and AI tool outputs.

Gibney offered several specific technical recommendations. First, models should receive only content that is actually intended for users. This means stripping out hidden HTML elements before passing email content to an AI summarizer. Second, organizations should implement controls for detecting attempts to conceal text through font size, color, or other formatting tricks. Third, email metadata should be clearly separated from message content when constructing prompts for the AI. The model should receive the email body in a way that makes it unambiguous which parts are user-facing content and which parts are structural or metadata.

Fourth, and perhaps most critically, AI-generated summaries should be checked against the original source. This could be done automatically by comparing key facts—numbers, dates, names—between the summary and the original email. If discrepancies are found, the summary should be flagged or discarded. Fifth, organizations should enforce a principle of least privilege on any actions an AI assistant can take. An AI that can read email should not automatically be able to send email, schedule meetings, or access payment systems. By limiting the AI’s authority, organizations can contain the damage even if a prompt injection attack succeeds.

Mapping the Attack Surface: The Hidden Infrastructure Problem

Gibney also highlighted a practical challenge that many organizations are only beginning to confront. “Organizations need to maintain an inventory of everywhere an LLM can read untrusted content for their users,” he said. “This is the attack surface, and in most organizations, it’s grown faster than has been mapped.”

The problem is that AI tools are often deployed ad hoc. A team might integrate an LLM-powered summarizer into their email client. Another team might use an AI tool to analyze customer support tickets. A third team might deploy an AI assistant for internal documentation. Each of these tools represents a potential entry point for prompt injection, and many organizations have no centralized visibility into where their AI systems are reading untrusted content. An attacker only needs to find one vulnerable integration to launch an attack.

The growth of this attack surface is accelerating. As AI tools become cheaper, easier to integrate, and more powerful, the number of places where LLMs interact with untrusted data will only increase. Organizations that do not map this surface proactively are effectively flying blind.

Industry Context: OWASP’s Persistent Warning

Forcepoint’s research validates warnings that have been issued at the highest levels of the security community. Since 2023, OWASP has ranked prompt injection as the number one risk for LLM and generative AI applications. The ranking appears in every edition of the OWASP Top 10 for LLM Applications, a document that has become the de facto standard for AI security risk assessment.

Other risks on the list include sensitive information disclosure, supply chain vulnerabilities, and excessive agency—all of which can be triggered or amplified by prompt injection attacks. The consistent top ranking reflects a consensus that prompt injection is both the most likely and the most damaging vulnerability that AI applications face. Forcepoint’s email summarization study is a concrete example of why that ranking is justified.

The research also connects to a broader category of attacks known as indirect prompt injection, where the malicious prompt is embedded in content that the AI retrieves from an external source. Email is one such source. Websites, documents, and APIs are others. Any system that processes untrusted external content is potentially vulnerable. The email vector is particularly concerning because email remains the backbone of corporate communication and because the hidden text technique is so easy to execute.

What Organizations Should Do Now

Gibney’s advice is direct: “Organizations need to secure their use of LLMs structurally; separate trusted instructions to a model from all untrusted content. They should also treat model output as untrusted and verify all model output back to the source text received by the model.”

Practically, this means security teams should take several immediate steps. Inventory all AI tools that process external content, especially email. Review the pipeline between the email source and the AI model to identify where hidden content could slip through. Implement HTML sanitization to strip invisible elements before text reaches the model. Establish a verification layer that checks AI summaries against original source material. And apply least-privilege controls to limit what AI systems can do with the information they process.

For organizations that have already deployed AI-powered email summarization tools, the risk is not hypothetical. Every email that passes through the system is a potential vector for attack. The hidden prompt technique is simple enough that any motivated attacker could implement it. The only question is whether organizations will act before the first successful attack, or after.

The Measurement That Changes the Conversation

Forcepoint X-Labs has done the security industry a service by converting a known vulnerability into a measured, repeatable result. The 100 percent success rate across ten trials provides a data point that security teams can use to make the case for investment in AI security controls. It is one thing to say that prompt injection is theoretically possible. It is another to show that, under realistic conditions, it works every single time.

The study also underscores the urgency of the agentic AI transition. As AI systems gain the ability to act on the information they process, the consequences of prompt injection attacks will escalate from misinformation to direct financial and operational harm. An AI that can send a fraudulent invoice is more dangerous than an AI that can only generate a false summary. Organizations that are building or deploying agentic AI systems must treat prompt injection as a primary design constraint, not an afterthought.

The hidden prompt problem is unlikely to be solved by better AI models. The fundamental architecture of LLMs makes them vulnerable to this class of attack, and no amount of training data or model tuning can fully eliminate the risk. The solution lies on the operational side: better input sanitization, output verification, and architectural separation of trusted instructions from untrusted content. Forcepoint’s research provides a clear roadmap for that work.

Share This Article