Claude Mythos Preview First to Clear All AISI Cyber Tests
Anthropic's latest Claude Mythos Preview checkpoint stands as the first AI model to complete every cyberattack simulation run by Britain's AI Safety Institute, or AISI. This model succeeded in a tough 32-stage assault on a mock corporate network during six of ten runs. It also broke through a simulated industrial control system known as Cooling Tower in three of ten attempts, a feat no prior model achieved.
AISI Updates Forecast on Rapid AI Cyber Gains
Frontier AI systems are building cyber skills quicker than forecasts predicted. The AISI first projected in November 2025 that these abilities would double every eight months. That estimate rose to 4.7 months by February 2026. Yet Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5 have gone far beyond even this faster pace, AISI reports. Experts question if this marks a lasting shift or just a sudden surge.
The AISI cyber ranges test models with detailed hacking scenarios. The corporate network challenge mimics a 32-step process that takes human specialists around 20 hours. The new Mythos Preview handled it fully in six out of ten cases. An earlier Mythos version only managed three out of ten. This checkpoint went live for partners too. On another test called The Last Ones, top models like Mythos Preview and GPT-5.5-Cyber hit phase nine, full network control, in best runs.
"The direction of travel is clear: cyber capabilities are advancing rapidly, and recent models represent a meaningful step up from what came before," AISI stated. The group now develops tougher tests with defenses to match the pace.
XBOW Evaluation Highlights Code Analysis Power
Security company XBOW ran independent checks on Mythos Preview with ten specialists. They called it a major step forward, noting unprecedented precision per token in spotting flaws. Against Anthropic's Opus 4.6, it reduced missed vulnerabilities by 42 percent. That figure climbed to 55 percent with source code available.
Mythos Preview shines brightest at reviewing code. "This was the first instance of a theme that would surface again and again: Mythos Preview is impressive at writing code, but even more impressive at reading it," XBOW noted. It even spotted issues in Chromium's V8 sandbox, where others only flagged false alarms.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Limits appeared too. Real vulnerabilities often stem from setups, dependencies, or component interactions, not just code. Even in pure code benchmarks, losing live system access hurt scores more than losing code access. The model excels at code review but needs system interaction for peak results.
Costs Temper the Model's Edge
XBOW questions if the gains justify rising prices. Anthropic set Mythos Preview at five times the cost of an Opus model. When adjusted for operating expenses, it performs well for accuracy but not top among benchmarks. Giving a GPT-5.5 agent extra time often matches or beats it cheaper. "The better option depends on the use case; often, it's the latter," XBOW said. They suggest using multiple models together.
Mythos Preview leads in straight vulnerability hunts like web apps or V8 sandbox. It rates mediocre or fair on complex judgment tasks.
Logan Graham, Anthropic's red-teaming lead for Project Glasswing, provided context. Partners using Mythos Preview found thousands of high and critical vulnerabilities in weeks, sometimes twice their yearly norm. Still, he cautioned, "Within a year, Mythos will probably look quite dumb (relative to other new models)." The key is readying for AI that outpaces top human experts in dual-use skills. Others might release similar open models.
Geopolitical Tensions Rise Over Access
This pushes action in software and policy. The US government reviews and tests Mythos. Anthropic bars China and the EU. OpenAI contacted the EU about GPT-5.5-Cyber access. Europe relies heavily on US firms, lacking homegrown rivals.
Anthropic, founded in 2021 by ex-OpenAI staff, prioritizes safe AI development. Its Claude series emphasizes alignment. The AISI, part of the UK government, evaluates frontier models for risks like cyber threats.
Related on Neura Market

