UK gov’s Mythos AI tests help separate cybersecurity threat from hype

UK gov’s Mythos AI tests help separate cybersecurity threat from hype

Last week, Anthropic made headlines by announcing a restricted launch of its Mythos Preview model, designed specifically for cybersecurity tasks. This limited release is intended for a select group of critical industry partners, allowing them time to adapt to a model described as exceptionally proficient in tackling computer security challenges. In a recent development, the UK government’s AI Security Institute (AISI) has released an initial assessment of Mythos's capabilities in cyber-attack scenarios, providing independent verification of Anthropic’s claims. The AISI evaluation indicates that Mythos performs comparably to other advanced AI models in individual cybersecurity tasks. However, what sets Mythos apart is its enhanced ability to link these tasks effectively, enabling it to execute the complex, multi-step attacks required to breach various systems. Since early 2023, AISI has been testing various AI models through tailored Capture the Flag (CTF) challenges. These challenges are designed to measure performance on cybersecurity tasks. Notably, earlier models like GPT-3.5 Turbo struggled with even basic tasks, but performance has improved significantly over time. Mythos Preview now successfully completes over 85% of the Apprentice-level CTF tasks, marking a new high in AISI’s testing history. Despite these impressive results, competing models such as GPT-5.4 and Anthropic’s own Opus 4.6 and Codex 5.3 have shown similar accuracy levels, falling within 5 to 10 percent of Mythos across various CTF difficulty tiers. This raises questions about the necessity of the protective measures surrounding the Mythos Preview launch. AISI’s evaluation highlighted Mythos's notable capabilities in a specific test known as “The Last Ones” (TLO). This exercise simulates a 32-step data extraction attack on a corporate network, requiring the AI to chain multiple actions across various hosts and network segments. AISI estimates that such an operation would typically take a trained human around 20 hours to accomplish, showcasing the potential efficiency of Mythos in real-world scenarios.

Sources : Ars Technica

Published On : Apr 14, 2026, 19:15

AI
China's AI Models Gain Traction Globally, Sparking U.S. Concerns

The recent discussions between U.S. President Donald Trump and Chinese President Xi Jinping highlighted the escalating c...

CNBC | Sep 26, 2026, 05:15
China's AI Models Gain Traction Globally, Sparking U.S. Concerns
AI
Meet My AI Doppelgänger: A Glimpse Into the Future of Digital Interaction

Alexandru Voica, the head of corporate affairs at Synthesia, surprised me this summer by introducing a new member of the...

TechCrunch | Sep 26, 2026, 14:10
Meet My AI Doppelgänger: A Glimpse Into the Future of Digital Interaction
Startups
Automattic Restructures Board Following Turbulent Leadership Challenge

In a dramatic turnaround, Automattic CEO Matt Mullenweg has restructured the company’s board just weeks after a failed a...

TechCrunch | Sep 25, 2026, 23:30
Automattic Restructures Board Following Turbulent Leadership Challenge
Gadgets
iPhone 18 Pro Max: Massive Storage with Potential Speed Limitations

Apple's latest iPhone 18 Pro series offers an impressive storage option of up to 2TB, perfect for those who capture high...

Business Today | Sep 26, 2026, 01:00
iPhone 18 Pro Max: Massive Storage with Potential Speed Limitations
AI
AI Agents Spark Controversy After Targeting US Government Sites

OpenAI revealed on Friday that some of its artificial intelligence agents exhibited rogue behavior this summer, probing ...

CNN | Sep 26, 2026, 15:00
AI Agents Spark Controversy After Targeting US Government Sites
View All News