T3MP3ST Framework Turns AI Coding Agents Into Autonomous 0-Day Hunters

An open-source framework repurposes AI coding agents like Claude Code and Codex for autonomous red-teaming and vulnerability discovery.

By Central
T3MP3ST achieves 90.1% pass@1 on XBEN benchmark and finds 8 of 10 real CVEs from 2026.
Highlights
  • T3MP3ST orchestrates AI coding agents through a full reconnaissance-to-exploit kill chain without requiring new API keys or cloud costs.
  • The framework scored 90.1% on XBOW's XBEN suite, outperforming XBOW's own self-reported 85% result.
  • A single agent found 8 of 10 real CVEs from 2026, demonstrating genuine autonomous 0-day hunting beyond model training data.

A newly released open-source security framework, T3MP3ST, is transforming general-purpose AI coding agents like Claude Code, OpenAI’s Codex, and Hermes into autonomous red-teaming operators, all without the need for new API keys, cloud infrastructure, or additional billing. Developed by the security researcher known as elder-plinius, T3MP3ST functions as a multi-agent orchestration layer rather than shipping its own proprietary model. It coordinates multiple agent instances through a complete reconnaissance-to-exploit-to-report kill chain, effectively repurposing AI coding tools already present on a user’s machine.

What Is the T3MP3ST Framework and How Does It Work?

Users can direct T3MP3ST at an authorized target using either a web-based “War Room” interface or a command-line interface (CLI). The AI coding agent already running on the operator’s machine then becomes the operational brain, driving the entire security mission. The framework is described as enabling “keyless warfare,” as it leverages existing agent sessions instead of demanding separate provider keys. It also enforces egress-scope containment, ensuring that networked tools automatically refuse to interact with off-scope public hosts, a critical safety feature for authorized testing.

T3MP3ST Performance Benchmarks and Autonomous 0-Day Hunting

T3MP3ST claims a 90.1% pass@1 score on XBOW’s 104-challenge XBEN suite, a black-box benchmark. This notably outperforms XBOW’s own self-reported result of roughly 85%. Every solve is graded against a committed flag oracle, and a “verify-claims” command can recompute results on demand for full reproducibility. On Cybench, a 40-task academic benchmark, the framework’s single-agent ReAct loop achieved 23 out of 40 hint-free solves.

The most striking results come from a held-out set of 10 real CVEs disclosed in 2026 across seven programming languages. In this test, a single agent successfully identified 8 of the 10 vulnerabilities, pinpointing the exact file, line, and CWE classification. The broader tool pack surfaced all 10 results. The developers frame these results as directional given the small sample size, but they are particularly meaningful because the bugs postdate the model’s training cutoff, essentially ruling out the possibility of training data memorization and demonstrating a genuine capacity for autonomous 0-day hunting.

The T3MP3ST Kill Chain and Current Status

The framework’s design maps an 8-operator kill chain—Recon, Scanner, Exploiter, Infiltrator, Exfiltrator, Ghost, Coordinator, and Analyst—onto MITRE ATT&CK tactics and the Cyber Kill Chain. However, it is critical to note that only the recon engine and single-agent exploit loop are currently benchmarked and considered stable. These core components can be cloned from the project’s public GitHub repository.

The downstream operators, which are designed to run the same tool-backed reasoning loop as the recon module, remain classified as experimental. End-to-end coordinated-swarm exploitation has not yet been validated at scale.

Domain Status Breakdown

The framework’s capabilities vary significantly by domain, as outlined in the project’s current status:

  • Web apps (XBEN suite): Stable and benchmarked.
  • CTF challenges (Cybench): Stable and benchmarked.
  • Embedded/OT/robotics OSS: Pipeline stable, coordinated disclosure in progress.
  • Source code (white-box): Experimental, Python-only ingest.
  • Smart contracts (DeFi): Experimental, reproduction only.
  • Cloud, mobile, AD, binary RE: On the roadmap and in development.

Reaction and Context in the Cybersecurity Community

Security researchers on platforms like Reddit’s blueteamsec community have flagged the release as a notable development for autonomous red-teaming. The trend is part of a broader industry momentum toward AI-driven security tooling, following related developments such as Anthropic’s Mythos model. XBOW separately evaluated Mythos and found it substantially improved vulnerability-led generation and source-code security analysis, cutting false negatives by 42% in comparable exploit benchmarks.

What This Means for Security Professionals and Organizations

The introduction of T3MP3ST signals a significant shift in the offensive security landscape. By lowering the barrier to entry for sophisticated, AI-driven penetration testing, it democratizes capabilities that were previously the domain of teams with substantial budgets and specialized infrastructure. For security teams, this means the potential to augment their red-teaming efforts with a powerful, cost-effective tool that can automate repetitive tasks and uncover complex vulnerabilities. However, it also means that defenders must be aware that adversaries may soon have access to similar, if not identical, autonomous capabilities. This underscores the urgent need for robust, layered defenses, including comprehensive patch management, network segmentation, and proactive threat hunting.

The developers explicitly state that T3MP3ST is strictly for authorized testing, research, and education. It is released under the AGPL-3.0 license with no warranty. Unauthorized use against systems without explicit written permission remains illegal in most jurisdictions. The responsibility for staying within legal and rules-of-engagement boundaries rests entirely with the operator. This is not a tool for casual experimentation or unapproved use.

Actionable Recommendations for Security Teams

For security teams looking to leverage this new capability, the immediate action is to assess your current red-teaming and vulnerability assessment workflows. Determine if and how an autonomous framework like T3MP3ST could be integrated into your authorized testing programs to improve efficiency and coverage. Simultaneously, reassess your organization’s defensive posture in light of the broader availability of AI-powered offensive tools. This should include accelerating patch cycles for known vulnerabilities, deploying advanced endpoint detection and response (EDR) solutions that use behavioral analysis rather than just signature-based detection, and conducting tabletop exercises that simulate attacks launched by autonomous agents. The goal is not just to use new tools, but to understand how your organization would fare against an adversary equipped with them.

Share This Article