Users in Houthi-Held Yemen Used Claude AI to Develop Weapons

Anthropic's report reveals that Houthi users attempted to use Claude AI for missile guidance, highlighting the new frontier of AI misuse in conflict zones.

By Central
Highlights
  • Anthropic's threat intelligence team identified a cell in northern Yemen attempting to develop missile guidance systems using Claude AI.
  • The Houthi users managed a failed rocket test and returned to Claude AI to analyze the failure, creating a dangerous feedback loop.
  • This case marks the first publicly known attempt by a non-state actor to use consumer-grade AI for weapons development.

When a company like Anthropic reports that users in Houthi-controlled Yemen attempted to use its Claude AI to develop advanced missile guidance systems, the implications extend far beyond a single blocked account. It crystallizes a new and dangerous reality: artificial intelligence, once confined to data centers and consumer applications, has become a viable tool for non-state actors seeking to modernize their arsenals. The report, the third from Anthropic since March 2025 on global AI misuse, reveals that a cell in northern Yemen pursued three distinct weapons programs, including a multi-variant hypersonic glide missile and a maneuverable warhead designed around mobile phone hardware. While the company states the users did not succeed in fielding an operational device, they did manage a failed rocket test — and crucially, they returned to Claude to understand why it failed. That feedback loop, between physical failure and AI-aided analysis, marks a threshold in how conflict zones are evolving.

Anthropic’s findings, covering the period from December to August, describe an operation that bypassed conventional software engineering. Instead of hiring human developers, the actors used the Claude Code interface to craft guidance, navigation, and control software for their projects. The goal, according to Anthropic’s report, was to develop a missile architecture that could accept multiple warhead and guidance configurations — different variants fitted to the same base design. This is a standard modern engineering approach, but for a group operating in a rugged, blockaded region, it represents an advanced logistical strategy. It suggests the Houthis aim to build a family of weapons from a single platform, reducing dependency on skilled labor and foreign components.

Anthropic did not identify the specific users, but mountainous northern Yemen is firmly controlled by the Houthis, the Iran-backed group officially known as Ansar Allah. This context places the attempt squarely within a broader pattern: the rebels are already wielding drones, conventional ballistic missiles, cruise missiles, and anti-ship munitions in their war against the Saudi-backed government and in attacks on shipping in the Red Sea. The campaign to seize strategic territory along the Red Sea coast, including the Bab al-Mandeb Strait, underscores their ambition. The attempt to use large language models like Claude indicates that the group’s technical leadership recognizes the potential of AI to compress the weapon development timeline.

Hazam al-Assad, a member of the political bureau of the Houthis, rejected the implication that his forces rely on any public tool for weapon making. He claimed their armed forces possess “modern, diverse and developed production capabilities” accumulated over years of fighting against Saudi-led Coalition. While this position is politically necessary for the group, it conflicts with the forensic evidence presented by Anthropic. The company says its threat intelligence team identified the cell, blocked the accounts, and shared the findings with both private and public partners. Anthropic claims that before the ban, the users had already built an offline simulation toolkit that functions independently of Claude or any external platform, indicating a concerning level of engineering momentum.

The Houthis Sought Hypersonic Capabilities Through Conversational AI

The details in the report paint a specific picture of what the users were trying to build. One of the projects involved a multi-variant missile intended to glide at hypersonic speed — exceeding Mach 5, the threshold for a dedicated program that even the United States has yet to fully operationalize. Trevor Ball, a weapons analyst at Armament Research Services, observed that while the group likely looked into hypersonic technology by asking the model, they lack the production capability to build such weapons currently. Representatives note that American hypersonic missiles remain in test phases, so the implication for a blockaded militia is even more severe. What the users could realistically accomplish, Ball noted, is the adaptation of existing Iranian anti-ship missiles — which already possess mid-course guidance adjustment — to reduce their dependence on Tehran shipments.

This observation is crucial because it frames the Claude incident as a path toward strategy, not just innovation. Yemen under blockade cannot rely on a steady supply of Iranian components. The United Nations embargo prohibits such transfers, and multiple independent inspections across the Arabian Sea and Gulf of Aden have seized Iranian-manufactured weapons contraband destined for Houthi forces. By using Claude to prototype guidance code, the Houthis could theoretically try to reverse-engineer or adapt existing parts into a homemade variant. This model of development is consistent with how battlefield innovation has occurred in other sanctioned environments — from improvised explosive devices in Iraq to missile work in North Korea. The difference here is the speed at which code can be generated and the availability of the platform.

Anthropic’s report describes that the users already had a failed guided rocket test. They returned to the Claude chatbot to diagnose it. This is a process human engineers often improve hardware by running failure analysis. For a non-state actor, this is an unprecedented degree of iterative design that previously would have required a team of physicists and aerospace engineers. The mere fact that the users tried this demonstrates that conversational AI is becoming a tool for tactical research, not just strategic propaganda, as previously seen in automated disinformation campaigns.

The Broader Global Pattern of AI-Arms Misuse

The Yemen disclosure is not an isolated case. Anthropic’s report, which covers activity from December to August, also documents state-sponsored groups spreading propaganda and unnamed actors using the Claude platform to research how to make biological weapons more deadly. This illustrates that AI providers are facing a dual-use dilemma far beyond text generation. The Houthi attempt fits into a category of misuse that Anthropic calls “mid-tier” capability build-up: actors who are not just asking prompts but actively using the AI to write code for operational military hardware.

For commercial AI developers, this forces them to serve as gatekeepers on par with encryption behind strong export controls. Anthropic has already begun implementing what it calls “trust and safety” layers specifically designed to detect when a user is trying to create harm. Their report explains that Claude has guardrails against generating weapons guidance systems and that these guardrails were triggered, leading to account removal. Yet the existence of a problem persists: the users remained capable of extracting enough learning to build the offline simulation tool before their ban. This indicates that even a short window of interaction with a frontier model can be enough to transfer dangerous knowledge.

This is relevant to how other companies also handle similar threats. Earlier reports from security researchers at the Center for Security and Emerging Technology (CSET) and from OpenAI itself had already warned that AI could lower the “capability threshold” for advanced weapons or biological weapons design. The fact that Anthropic contributed its own analysis adds weight to the idea that accidents happen in war.

How the Houthis Might Be Using AI for Self-Sufficiency

To understand the significance: the Houthi forces have always relied on a combination of Iranian components, indigenous workshop-level production, and foreign technical assistance. For years, they could not develop truly homegrown high-end systems. The use of a large language model like Claude could theoretically allow them to produce the guidance code for a missile variant without needing to import a team of engineers or create a university-level computer science curriculum. In effect, they can use the AI to “translate” complex aerospace problems into small, practical scripts that a well-trained local technician can implement.

The Houthi’s rise is especially notable given that they increasingly threaten global trade through the Bab al-Mandeb Strait. The blockage of that route, a gateway connecting the red sea with the Gulf of Aden, costs global shipping billions each week in rerouting. The fact that the group involved in the weapons testing is itself the same group that attacks tankers and oil facilities suggests that the usage of AI is being driven by immediate tactical goals, not abstract experimentation. These attacks to pressure Saudi Arabia and help disrupt global supply chains in conjunction with other Axis of Resistance campaigns in the region.

Adam Baron, a Yemen-focused researcher at the New America think tank, summarized that the Houthis have kept pace with technological developments. He warned that there is a tendency to perceive them as “barefoot tribal fighters,” but that is irrelevant. Their strategic use of social media narratives, their ability to exploit transfers of Iranian expertise, and now their tentative use of AI — everything points to an “increasing amount of institutional tech specialty.” This matches everything we have reported on the group’s progress over the last decade.

From an operational standpoint, the Houthi’s ability to carry out even a failed test based on the code generated from a conversation with an AI highlights a gap in the AI arms control regime. Traditional regulations, like the Wassenaar Arrangement on export controls for conventional weapons and dual-use goods, cover physical hardware for guided weapon aboard. They do not cover a general-purpose AI model accessible from a laptop. As the ability to generate offline improves, the barrier to entry for acquiring near-peer knowledge falls. This is the fundamental shift: not that the Houthis built hypersonic, but that they attempted to build code for them.

Strategic Implications for Defense and AI Policy

The incident raises two separate but overlapping: one regarding defense readiness, the other regarding AI safety. From a defensive vantage point, the Pentagon and its allies already prioritize AI for countering drone threats and intelligence fusion. What this disclosure suggests is that AI is also a growing offensive equalizer for non-state actors. The confirmation from Anthropic that the users developed an offline tool kit means that even if the model is banned, the knowledge remains in the offline environment. This “export of the ability” cannot be turned off by terminating an account. It is already lived in the minds and documents of the users.

To address this, policymakers need to reassess the scope of dual-use controls to include not just knowledge but the tools that generate knowledge. This might require AI vendors to implement dynamic detection based on geographic context or specialized domain modelling, such as detecting patterns that indicate missile flight dynamics or guidance loop physics. The report’s mention of a “shut down offline simulation tool kit” suggests that the early detection or intervention may also leave behind residual knowledge.

At the same time, Anthropic’s decision to publish this information shows that transparency is also a tool for accountability. By sharing these red flags, the company pressures other labs to adopt similar standards of monitoring and disclosure. The specific pattern of asking Claude to generate a guidance system would be a different issue than general queries. For example, a query that matches a known guidance algorithm should phrase the output to avoid providing a information ready that allows assembly.

In the longer term, the matter shows that AI cannot remain neutral. The very same AI that can be used to write code for a humanitarian drone delivery could be re-purposed, within minutes, by a different group, to write control code for a warlike intercept system. This demands that LinkedIn and other executives not only set boundaries but also enforce them in the context of geopolitical conflict.

The Houthi example signals that the future of warfare will not be about access to the most secret supercomputers. Instead, it will be about how best to utilize a civilian-available, consumer-grade AI to accelerate the gap between intent and capability. For Yemen and many other conflict zones, the AI revolution is already being tested on the battlefield as a cost-effective force multiplier, and this particular case is merely the first known attempt that has been made public.

Share This Article