Anthropic made a major announcement on April 7, 2026, unveiling Project Glasswing. This project offered select organizations early access to their new Claude Mythos Preview model. The purpose of this initiative was to allow cyber defenders to identify and address critical software vulnerabilities before the model was released to the public. While some skeptics questioned the validity of Glasswing, recent reports from Anthropic indicate that participants have already uncovered 10,000 high or critical severity vulnerabilities. Notable organizations like Cloudflare, Mozilla, and Palo Alto Networks have been successful in finding significant vulnerabilities. However, the debate continues, with some emphasizing that discovering a vulnerability is only the first step, and exploiting it is a much more complex challenge.
The effectiveness of Mythos and other cutting-edge models in exploitation is a crucial question that the new AI vulnerability exploitation benchmark, ExploitGym, seeks to answer. The results so far indicate that recent frontier models are highly proficient, fast, and resourceful, hinting at a future where AI can autonomously discover and exploit vulnerabilities at remarkable speeds.
Despite the impressive capabilities of these models, the exploits and tactics utilized are based on familiar principles. While standard defenses have shown decent performance in mitigating attacks, they are not sufficient on their own. An effective defense-in-depth strategy, starting with enhanced model safety, is essential in combating evolving cyber threats. Let’s delve deeper into these findings by exploring the ExploitGym benchmark.
Understanding ExploitGym
ExploitGym, developed by a team of researchers from various institutions including UC Berkeley, Max Planck Institute for Security and Privacy, and Google, is a comprehensive benchmark designed to evaluate the vulnerability exploitation capabilities of state-of-the-art LLMs. These models, released between February and April, are renowned for their prowess in coding and cybersecurity.
The benchmark comprises 898 unique instances, each consisting of a real-world vulnerability that has impacted widely-used software projects in three key domains: Userspace Programs, The Browser (specifically Google V8 JavaScript engine), and The Kernel (Linux kernel). The objective for AI agents in each instance is to achieve unauthorized code execution within a strict two-hour timeframe, simulating a high-impact security breach scenario.
To ensure fair evaluation, a verified-access framework was employed, relaxing certain model safeguards to push the boundaries of their capabilities. The results shed light on the evolving landscape of AI-driven vulnerability exploitation and the efficacy of standard defenses against such attacks.
Key Takeaways from the Results
The findings from ExploitGym underscore several important points:
1. Recent Models Surpass Human Exploitation Capabilities
While the numbers may seem unimpressive at first glance, the success rates achieved by leading models without defenses activated are remarkable. In challenging scenarios where time constraints and difficulty levels are high, these models outperform human experts in vulnerability exploitation. This highlights the potential of AI in enhancing cybersecurity efforts.
2. Standard Defenses Offer Significant Protection
The effectiveness of standard defenses is evident in the substantial reduction of success rates when activated. However, the ability of models like Claude Mythos Preview and GPT-5.5 to bypass these defenses in a significant number of instances underscores the need for continuous evolution in defense strategies.
3. Evolution within Model Families
The performance gap between Claude Mythos Preview and other models within the same family, as well as the notable success of GPT-5.5, indicates a verified step change in exploitation capabilities. These advancements hint at a future where Mythos-level models will become more prevalent in the cybersecurity landscape.
4. Rise of Fully Automated Exploitation
A significant portion of initial successes in exploitation was achieved through ‘cheating,’ demonstrating the autonomy and resourcefulness of AI agents. These findings suggest a shift towards fully automated exploitation methods, where AI can adapt and pivot strategies based on evolving circumstances.
5. Role of Model-Based Mitigations in Defense
The results highlight the importance of model-based mitigations in a defense-in-depth strategy. While standard defenses provide a foundational barrier, models play a crucial role in enhancing overall security posture. By leveraging insights from exploitation research, organizations can strengthen their defenses and stay ahead of emerging cyber threats.
Exploitation research serves as a valuable tool in improving cybersecurity practices, enabling better understanding and mitigation of AI-driven attacks. By staying informed and proactive in addressing vulnerabilities, organizations can bolster their defenses and uphold the integrity of their systems amidst evolving threats.
Laura Wilber, Senior Analyst at Enea AB, specializes in AI safety and security, post-quantum computing, and critical network defense. With a focus on research and analysis, Laura supports teams in navigating complex cybersecurity challenges across various domains.
For more information, you can connect with Laura on LinkedIn or visit the Enea website.

