Prompt Injection Ranks No. 1 with OWASP, No. 12 in Incidents

A new study reveals a stark disconnect between expert fears and real-world LLM security incidents.

By Central
Prompt injection ranks first in OWASP expert surveys but only 12th in actual incident data.
Highlights
  • Prompt injection has held the No. 1 spot on OWASP's Top 10 for LLMs for three consecutive years.
  • The study found no statistically detectable agreement between expert rankings and incident data.
  • Prompt injection operates in a part of the system where vulnerability scanners cannot detect it.

The disconnect between what security experts fear most about large language models and what the public incident record actually shows has never been starker. Prompt injection has held the No. 1 spot on the OWASP Top 10 for LLM Applications for three consecutive years, yet when two leaders of that project cross-referenced the expert ranking against 6,639 labeled real-world incidents, the attack landed at No. 12. The disparity is not a measure of reduced danger. It is a measure of visibility, because prompt injection operates in a part of the system where vulnerability scanners cannot see it, and the gap between what experts warn about and what gets reported has profound implications for how enterprises budget AI security.

That finding comes from Kyriakos “Rock” Lambros and Steve Wilson, two leaders of the OWASP Top 10 for LLM Applications project, who published their analysis on arXiv on August 18. The paper carries a disclaimer: it is exploratory, not peer reviewed, and not the official OWASP release. The disclaimer matters, but the machinery behind it is real. The researchers assembled 7,714 LLM security incidents drawn from CVE, GitHub Security Advisories, OSV, and the AIAAIC AI-harm database. They labeled 6,639 of those incidents against a 20-entry taxonomy and used a Bayesian model to correct each count for classifier error, then set the data-driven ranking beside the expert vote.

The comparison found no statistically detectable agreement between the two rankings. Cohen’s kappa, a standard measure of inter-rater reliability, came in at 0.20 with a 90% confidence interval running from negative 0.16 to 0.57. “The interval crosses zero, so we cannot rule out that the two rankings agree only by chance,” the authors write. “The honest bottom line: weak agreement, not confirmation.” Lambros, co-lead of the OWASP GenAI Security Project Top 10 for LLM Applications and director of AI standards and governance at Zenity, put the finding in stark evidentiary terms. “We had two ways of measuring the same risk, expert judgment and the public incident record, and they disagree with each other. Neither one is the truth. Two witnesses are contradicting each other, and we can’t tell you which one is lying.”

The attack chain a scanner never logs

The structural reason for the gap is that prompt injection does not leave a product defect for a scanner to find. The attack hides instructions inside the content a model reads — a log entry, a support ticket, a document pulled back by retrieval. An agent then makes the tool call the attacker wanted, using credentials it legitimately holds. Nothing in that chain is a bug. No CVE is generated. A vulnerability scanner looking for known weaknesses will pass right over a system that is being actively exploited through its own content ingestion pipeline.

The defenses that catch prompt injection are not scanner-based. They are adversarial tests run against the deployed system and hard caps on what the agent can reach, so a fooled model cannot touch anything expensive. The logic argues for funding agent memory and MCP tool boundaries now, on architecture, rather than waiting for advisory volume that will always arrive a cycle late.

The first control Wilson would deploy against an injected agent

Wilson, Chief AI and Product Officer at Exabeam and project co-lead for the OWASP Top 10 for LLM Applications, named the control he would deploy first against the exact attack chain — an agent that reads an attacker’s payload in a log file, treats it as an instruction, and rewrites DNS with a valid credential. “The first thing I’d do is put an authorization gate outside the model: the agent can propose the exact DNS change, but it cannot grant itself the authority to make it,” Wilson said. “Security rules written inside prompts may shape the model’s behavior, but they are still suggestions to the model, not enforceable security controls.” The tradeoff is explicit. The agent loses the ability to improvise arbitrary, high-impact infrastructure changes on its own. It retains autonomous investigation and routine, bounded remediation.

What is the difference between expert risk ranking and incident data for LLM threats?

The expert risk ranking is a survey of practitioner opinion. The incident data is a record of what has been observed, recognized, classified, and reported. The two measure fundamentally different things. The expert ranking captures what a small group of knowledgeable people believes is most dangerous based on their understanding of system architecture and attack surfaces. The incident data captures what has actually been caught and filed. Neither is truth. The expert ranking may overestimate threats that are well-understood but already defended against. The incident data may miss entire categories of attack that leave no reportable trace, such as memory poisoning or prompt injection that succeeds without the victim ever knowing. The honest answer is that security teams need both signals, and they need to understand that the two will often disagree.

Why the No. 1 risk looks small in the public record

The authors compress the entire divergence into a single sentence. “Experts rank it first because the attack surface stays enormous even when the defenses mostly hold; the data sees the successes that got through.” Prompt injection is the best-understood LLM attack. Deployed systems defend against it actively. A control that works most of the time still leaves an enormous attack surface, because the model is being asked to interpret trusted instructions and untrusted content simultaneously. A low advisory count can mean the defenses are working. It can just as easily mean nobody has looked. The public record cannot tell a security team which one it is.

Wilson has watched the gap from both sides. “Incident data is incredibly valuable, but it is inherently backward-looking and notoriously tricky to interpret. It tells us what was observed, recognized, classified, and reported. It does not necessarily tell us what is most dangerous in the systems people are building right now.” He compares prompt injection to a law of physics for LLM systems. A control that works 99% of the time is not sufficient when the failure case gives an attacker meaningful access. “The durable answer is not believing we can perfectly screen prompt injection out of existence. It is designing systems with the assumption that prompt injection will occur, understanding why it works, and limiting what an attacker can accomplish when it does.”

The attempt volume is documented. CrowdStrike’s 2026 Global Threat Report found adversaries injected malicious prompts into legitimate GenAI tools at more than 90 organizations in 2025, stealing credentials and cryptocurrency, under a section titled “Prompts are the New Malware.” That telemetry shows pressure on the attack surface without proving that defenses produced the No. 12 placement. It is the pattern the structural mechanism predicts.

The gap runs the other way too, and further for misinformation

Prompt injection is the headline case, but misinformation is the bigger divergence. The expert vote puts misinformation at No. 13. The incident record places it at No. 2. The paper calls it “the widest disagreement between the two witnesses” and reports that its concordance flag puts the probability that the two signals disagree at 99 percent. The authors do not treat their own data as the winner. On misinformation they note the corpus carries a large volume of deepfake and AI-generated disinformation, records that often describe harm produced by an AI rather than a vulnerability inside an LLM. The entry is the one the record most disputes, and the authors stop short of concluding the experts got it wrong.

Where “too new to measure” runs into the CVE record

The two brand-new taxonomy entries sit at the sharpest end of the gap. Persistent memory poisoning lands at expert No. 4 and incident No. 16. MCP tool interface exploitation lands at expert No. 7 and incident No. 16. Each carries an incident interval of 6 to 20 that spans most of the taxonomy, meaning the model cannot place either entry within 14 rank positions. Public 2026 CVEs exist for both. On MCP tool interfaces, the Azure Data Explorer MCP Server carried KQL injection, described in the CVE record as allowing “an attacker (or a prompt-injected AI agent) to execute arbitrary KQL queries against the Azure Data Explorer cluster,” scored 8.3 High. Kong’s Konnect MCP Server shipped an indirect prompt injection that lets a remote attacker steer the server into executing unintended API requests, the exact failure the MCP entry names.

Agent memory has its own record. An agent harness called Ruflo exposed unauthenticated MCP bridge endpoints that let a network attacker obtain a shell, read provider API keys, and poison the learning store. It was rated 10.0 Critical. A team waiting for advisory volume to justify a control on agent memory or an MCP tool boundary would still be waiting while the CVEs accumulate at High and Critical. Lambros makes the budget case in operational terms. Poisoned memory “doesn’t announce itself. Nobody files an advisory for that, because nobody knows it happened. A count of zero is measuring your blindness, not your safety.” The argument he says a CFO will sign off on is timing. Memory and tool permissions get wired into these systems once, early. Everything else sits on top of them. “Build it in now and it’s a rounding error. Come back in two years and you’re re-architecting and re-training your systems.”

The authors flag their own measurement problems first

The expert side is thin. “The expert signal is a practitioner survey: about 29 respondents scored each candidate risk on importance,” the authors write. Twenty-nine votes set the ranking that carries three-quarters of the published list’s weight, the compression point for OWASP’s more than 25,000 community members. On the data side, the classifier is the weak joint. Precision varies sharply across entries, from 93% for prompt injection and supply chain down to 13% for vector and embedding weaknesses. Four entries fall below 50% precision. The base classifier never predicts “out of scope” and files every incident into some category, including the roughly 38% of the gold set that belongs in none.

The authors name the central limitation themselves. One reviewer adjudicated all 1,200 gold-set incidents and overrode the model consensus on 553 of them. “A single annotator cannot measure inter-rater reliability. The single-author gold set remains the central limitation.” Lambros lays the weak kappa at the feet of the taxonomy itself. “That number is telling you about our categories, not about our experts. A weak score on the ordering of those buckets is a fact about the buckets.”

A better classifier will not fix the disagreement. A pre-registered bake-off of four frontier models produced no winner. None beat the incidence floor’s balanced accuracy of 0.863. A ground-truth check left the floor’s ordering in place at a Spearman correlation of 0.918. The authors published the engine and artifacts on GitHub for anyone to rerun. The robustness result tested only one side of the gap. Every check behind the word “robust” runs on the incident side, showing the incident-derived ranking stays put when the labeling machinery changes. None of it touches the 29-vote survey. A board that hears “robust” will assume validated. The record supports only stable.

What the published OWASP list did with this data

OWASP shipped the GenAI LLM Top 10 2026 on August 4, the first edition to fold incident data into the ranking. The weighting was 75% practitioner vote and 25% incident corpus. Prompt injection stayed at No. 1. Misinformation moved up two places. Excessive agency climbed from No. 6 to No. 3 as the entry where the two signals agree most clearly. Unbounded consumption rose four spots to No. 6. Improper output handling fell from No. 5 to No. 10, the largest drop. Wilson declines to defend the blend as arithmetic. “There is nothing magical about a 75/25 weighting, or about reversing it to 25/75. The value of the data wasn’t that it gave us a mathematical answer; it changed the conversation.” The excessive agency entry is where that conversation landed hardest for him. “If I were a CISO evaluating a new agentic deployment today, Excessive Agency is where I would start.”

Lambros would go further next cycle, a view he flags as his own and separate from the working group. The blend hands the same 25% incident weight to every category, while the hand-checked classifier precision runs from roughly nine in ten on prompt injection and supply chain down to roughly one in eight on vector and embedding weaknesses. A quarter of the weight on the first rides on something solid. The same quarter on the second rides on noise. “The ratio should track how well we actually measure each category.”

Why this lands now: 87% of security teams are prioritizing agentic AI

Ivanti’s 2026 State of Cybersecurity research found 87% of security teams call adopting agentic AI a priority and 77% report at least some comfort letting AI act without human review. Teams are signing off on agent autonomy while the expert ranking of what can go wrong with those agents shows no statistically detectable agreement with the incident record. The timing matters because the decisions being made today — which controls to fund, which risks to accept, which systems to deploy — are being made against a risk picture that the data cannot confirm.

What to do with this on Monday: five practical actions

The behavioral change is narrow and it is the whole point.

Use the OWASP LLM Top 10 as a coverage map, not a queue. The rank positions carry 29 votes and a corpus whose own authors call the agreement weak. Build your own priority order from your own exposure: production reach, breach-notification data, and controls that have actually been tested. Lambros draws the funding line the same way. “I’d prioritize spend where the expert vote and the incident record point the same direction, because that’s two independent witnesses agreeing. Where they split, stop letting the ranking allocate your money and go look at what your own systems are doing.”

Log what your AI systems are actually doing, field by field. The prompt that went in, what came back out, the documents pulled to build the answer, the tools called and the arguments passed to them, and the model’s confidence score on every response. Confidence is the field Lambros would fight for, because most security leaders do not realize it is measurable, and it is where the attack surfaces. “A model running on a poisoned instruction doesn’t act broken. It acts certain. Certainty is what your monitoring treats as a healthy system.” The cost is a sprint or two of engineering. The constraint is a person, because a SIEM does events and these are trends. “Somebody has to analyze those trends every week and say whether a drift means anything, and most security teams have nobody who can.”

Stop expecting scanner output to reproduce the Top 10’s order. Scanner findings live on the incident side of the gap, counting what got disclosed rather than what a deployed system should fear. The classifier bake-off shows a smarter model does not close that distance. The test that sees prompt injection is an adversarial one run against the live system, paired with Wilson’s authorization gate so the change an injected agent proposes is never the change it can execute.

Fund the thin-record categories on architecture, not incident volume. Agent memory and MCP tool boundaries sit at expert No. 4 and No. 7 with incident intervals spanning most of the taxonomy. The CVEs that do exist are landing at High and Critical. Kayne McGladrey, an IEEE senior member who advises enterprises on risk, put the funding logic bluntly. “Anything that seems to have a cybersecurity flavor is generally put into the cybersecurity risk category, which is a complete fiction. They should be focused on business risks, because if it doesn’t affect the business, like a financial loss, then nobody’s going to pay attention to it, and they will not budget it appropriately.” A rank number from a 29-person vote is a weaker budget argument than the business system the agent touches.

Steal McGladrey’s baseline test for the AI systems themselves. “If you wouldn’t expose your database to the public internet without identity and access controls, why would you do that for your AI model?”

The board question for the next meeting is short. If the AI risk ranking came from a 29-person vote and a corpus that disagrees with it, what are we actually using to decide which controls get funded next year? The answer has to be better than “the list.” It has to be the system itself.

Share This Article