AI Models

Claude Mythos Preview First to Clear All AISI Cyber Tests

Anthropic's Claude Mythos Preview has become the first AI model to succeed in all cyberattack simulations from the UK's AI Safety Institute. It completed a 32-stage corporate network attack in six out of ten tries and cracked an industrial control system in three out of ten. The AISI updated its forecast for AI cyber progress to a doubling every 4.7 months, which models like Mythos and GPT-5.5 have already beaten.

Neura News

Neura News

Neura Market Editorial

May 14, 20264 min read

Originally reported by the-decoder.com

Claude Mythos Preview First to Clear All AISI Cyber Tests

Claude Mythos Preview First to Clear All AISI Cyber Tests

Anthropic's latest Claude Mythos Preview checkpoint stands as the first AI model to complete every cyberattack simulation run by Britain's AI Safety Institute, or AISI. This model succeeded in a tough 32-stage assault on a mock corporate network during six of ten runs. It also broke through a simulated industrial control system known as Cooling Tower in three of ten attempts, a feat no prior model achieved.

AISI Updates Forecast on Rapid AI Cyber Gains

Frontier AI systems are building cyber skills quicker than forecasts predicted. The AISI first projected in November 2025 that these abilities would double every eight months. That estimate rose to 4.7 months by February 2026. Yet Anthropic's Claude Mythos Preview and OpenAI's GPT-5.5 have gone far beyond even this faster pace, AISI reports. Experts question if this marks a lasting shift or just a sudden surge.

The AISI cyber ranges test models with detailed hacking scenarios. The corporate network challenge mimics a 32-step process that takes human specialists around 20 hours. The new Mythos Preview handled it fully in six out of ten cases. An earlier Mythos version only managed three out of ten. This checkpoint went live for partners too. On another test called The Last Ones, top models like Mythos Preview and GPT-5.5-Cyber hit phase nine, full network control, in best runs.

"The direction of travel is clear: cyber capabilities are advancing rapidly, and recent models represent a meaningful step up from what came before," AISI stated. The group now develops tougher tests with defenses to match the pace.

XBOW Evaluation Highlights Code Analysis Power

Security company XBOW ran independent checks on Mythos Preview with ten specialists. They called it a major step forward, noting unprecedented precision per token in spotting flaws. Against Anthropic's Opus 4.6, it reduced missed vulnerabilities by 42 percent. That figure climbed to 55 percent with source code available.

Mythos Preview shines brightest at reviewing code. "This was the first instance of a theme that would surface again and again: Mythos Preview is impressive at writing code, but even more impressive at reading it," XBOW noted. It even spotted issues in Chromium's V8 sandbox, where others only flagged false alarms.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

Limits appeared too. Real vulnerabilities often stem from setups, dependencies, or component interactions, not just code. Even in pure code benchmarks, losing live system access hurt scores more than losing code access. The model excels at code review but needs system interaction for peak results.

Costs Temper the Model's Edge

XBOW questions if the gains justify rising prices. Anthropic set Mythos Preview at five times the cost of an Opus model. When adjusted for operating expenses, it performs well for accuracy but not top among benchmarks. Giving a GPT-5.5 agent extra time often matches or beats it cheaper. "The better option depends on the use case; often, it's the latter," XBOW said. They suggest using multiple models together.

Mythos Preview leads in straight vulnerability hunts like web apps or V8 sandbox. It rates mediocre or fair on complex judgment tasks.

Logan Graham, Anthropic's red-teaming lead for Project Glasswing, provided context. Partners using Mythos Preview found thousands of high and critical vulnerabilities in weeks, sometimes twice their yearly norm. Still, he cautioned, "Within a year, Mythos will probably look quite dumb (relative to other new models)." The key is readying for AI that outpaces top human experts in dual-use skills. Others might release similar open models.

Geopolitical Tensions Rise Over Access

This pushes action in software and policy. The US government reviews and tests Mythos. Anthropic bars China and the EU. OpenAI contacted the EU about GPT-5.5-Cyber access. Europe relies heavily on US firms, lacking homegrown rivals.

Anthropic, founded in 2021 by ex-OpenAI staff, prioritizes safe AI development. Its Claude series emphasizes alignment. The AISI, part of the UK government, evaluates frontier models for risks like cyber threats.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read