Agentjacking Tricks AI Coding Agents Into Running Malicious Code

Security researchers uncover Agentjacking, a new attack that tricks AI coding agents into executing malicious code through trusted error-reporting services.

By Central
Agentjacking exploits Sentry's public DSN and AI agents' implicit trust to execute arbitrary malicious payloads.
Highlights
  • Agentjacking lets attackers execute arbitrary code on developer machines without compromising infrastructure or phishing.​
  • The attack exploits Sentry's public Data Source Name (DSN) and the Model Context Protocol (MCP) used by AI coding agents.​
  • Sentry declined to fix the vulnerability, stating it is 'technically not defensible' at the platform level.

Security researchers have disclosed a novel attack technique, dubbed Agentjacking, that exploits the implicit trust AI coding agents place in external error-reporting services, enabling attackers to execute arbitrary code on developer machines without compromising infrastructure. The attack, uncovered by Tenet Security, weaponizes Sentry, a widely used open-source error-tracking and performance-monitoring platform, by injecting malicious payloads into error events that AI assistants such as Claude Code and Cursor subsequently interpret as legitimate remediation steps.

How Agentjacking Works: Abusing the Sentry Ecosystem

Agentjacking targets a fundamental architectural weakness at the intersection of Sentry’s event ingestion pipeline and the Model Context Protocol (MCP) used by AI coding agents. Sentry’s Data Source Name (DSN) is a public, write-only credential embedded in websites and applications. Because the DSN is designed to accept arbitrary payloads from anyone, an attacker can send a crafted error event to Sentry’s ingest endpoint via a simple POST request.

The injected event contains carefully formatted markdown in the message field and context key names. When the Sentry MCP server returns this event to an AI agent, the agent renders the markdown as structured content that is visually identical to Sentry’s own system template. A developer who then prompts their AI assistant to “fix unresolved Sentry issues” triggers the agent to query Sentry via MCP and retrieve the malicious event. The agent, unable to distinguish a legitimate error from an attacker-injected one, executes the embedded instructions with the developer’s full system privileges.

Successful exploitation can expose sensitive data including environment variables, Git credentials, private repository URLs, and developer identities. Critically, the attack chain bypasses every common security control—endpoint detection and response (EDR) systems, web application firewalls (WAFs), identity and access management (IAM) solutions, VPNs, Cloudflare protections, and firewalls—because every action in the chain is authorized from the system’s perspective.

What Is Agentjacking and Why Is It Dangerous?

Agentjacking is a class of attack that tricks AI coding agents into executing malicious code by injecting fake error reports into an error-tracking platform. It is dangerous because it exploits the trust developers place in their AI assistants and in tools like Sentry, using publicly available data (the DSN) as an entry point. No phishing, no prior server compromise, and no malware is required. The attack succeeds because AI agents currently lack the ability to distinguish between genuine system output and attacker-controlled content when both arrive through the same trusted channel.

The Attack Chain in Detail

Tenet Security outlined the complete attack chain in four steps:

  • An attacker identifies a target’s Sentry DSN, which is publicly embedded in websites and applications.
  • The attacker sends a malicious error event to Sentry’s ingest endpoint via a POST request, using the DSN as the authentication token.
  • The injected event contains markdown that, when rendered by the Sentry MCP server, appears identical to Sentry’s own system template. The message field and context key names are crafted to include attacker-controlled commands.
  • When a developer asks their AI coding agent to fix unresolved Sentry issues, the agent queries Sentry via MCP, retrieves the malicious event, and executes the embedded instructions with the developer’s full privileges.

The attacker never touches the victim’s infrastructure. The malicious instruction arrives disguised as a legitimate resolution inside an ordinary error.

Scale and Impact: 85% Success Rate Across Major AI Coding Assistants

Tenet Security identified at least 2,388 organizations with valid injectable DSNs exposed. In controlled testing against over 100 organizations, the researchers achieved an 85% exploitation success rate against injected errors across some of the most widely used AI coding assistants. The attack works against agents that rely on MCP to connect to external services, a category that includes Claude Code, Cursor, and similar tools increasingly adopted by development teams.

Sentry acknowledged the issue but declined to fix it, stating that the vulnerability is “technically not defensible” at the platform level. The company has, however, activated a global content filter designed to block a specific payload string associated with the demonstrated attack.

Broader Implications for Enterprise Security

Agentjacking represents a paradigm shift in the attack surface facing modern development environments. Traditional security controls assume that threats arrive as malicious files, phishing links, or unauthorized access attempts. This attack exploits the authorized data flows that developers themselves have configured, turning trusted tools against their users. As enterprises accelerate their adoption of AI coding agents, the trust boundary between an agent and the services it connects to becomes a critical vulnerability.

What Developers and Organizations Should Do Now

Until AI coding agents can independently verify the authenticity of external data sources, developers should treat all output from connected services with suspicion. Organizations should audit their exposure of Sentry DSNs and consider rotating any that are publicly accessible. Developers should review the prompts they use with AI coding agents and avoid open-ended instructions such as “fix all unresolved issues” without manual verification. Implementing a security policy that requires human approval before an AI agent executes any command that modifies files, accesses credentials, or interacts with production systems can reduce the risk. For the longer term, organizations should push for MCP implementations that include content validation and provenance checks, ensuring that AI agents can distinguish between genuine system output and attacker-injected data.

Share This Article