Amid Mythos’ hyped cybersecurity prowess, researchers find GPT-5.5 is just as good

Amid Mythos’ hyped cybersecurity prowess, researchers find GPT-5.5 is just as good

In a recent development, researchers from the UK’s AI Security Institute (AISI) have unveiled findings that challenge the claims made by Anthropic regarding its Mythos Preview model. After the company emphasized the cybersecurity threats posed by Mythos, it limited its initial release to select industry partners. However, the latest evaluations indicate that OpenAI’s recently launched GPT-5.5 exhibits comparable performance in cybersecurity tasks. Since the beginning of 2023, AISI has rigorously tested various AI models, including Mythos and GPT-5.5, through 95 different Capture the Flag challenges. These challenges assess capabilities in crucial areas such as reverse engineering, web exploitation, and cryptography. Remarkably, GPT-5.5 achieved an average success rate of 71.4% on the most challenging “Expert” tasks, outpacing Mythos Preview’s 68.6%, although this difference falls within the margin of error. One standout moment in the evaluation involved a particularly tough challenge that required creating a disassembler to decode a Rust binary. AISI reported that GPT-5.5 completed this task independently in just over 10 minutes, incurring an API cost of $1.73. Furthermore, in the AISI’s test scenario known as “The Last Ones” (TLO), which simulates a complex data extraction attack, GPT-5.5 succeeded in 3 out of 10 attempts. In contrast, Mythos managed just 2 out of 10 attempts, marking a significant achievement as no previous AI models had succeeded at this test before. Despite these accomplishments, GPT-5.5, like its predecessors, struggled with AISI’s more intricate “Cooling Tower” simulation that aims to disrupt power plant control software. This challenge remains a barrier that all tested AI models have yet to overcome.

Sources : Ars Technica

Published On : May 01, 2026, 15:35

Computing
Google Faces €890 Million Penalty in Landmark EU Digital Markets Act Enforcement

European regulators have imposed a hefty fine of €890 million (approximately $1 billion) on Google, citing the company's...

CNBC | Jul 23, 2026, 10:25
Google Faces €890 Million Penalty in Landmark EU Digital Markets Act Enforcement
AI
Elon Musk Embraces AI Risks for a Future of Abundance

In a thought-provoking interview with Zanny Minton Beddoes, the editor in chief of The Economist, Elon Musk shared his i...

Business Insider | Jul 23, 2026, 12:20
Elon Musk Embraces AI Risks for a Future of Abundance
AI
IMF Sounds Alarm on AI Risks in Finance, Urges Enhanced Oversight

The International Monetary Fund (IMF) has raised significant concerns regarding the impact of artificial intelligence on...

Business Today | Jul 23, 2026, 12:40
IMF Sounds Alarm on AI Risks in Finance, Urges Enhanced Oversight
AI
AI Trends Transforming India's IT Landscape: A Shift in Revenue Models

India's IT services sector, valued at approximately $300 billion, is on the brink of a significant transformation as art...

Business Today | Jul 23, 2026, 04:55
AI Trends Transforming India's IT Landscape: A Shift in Revenue Models
AI
Moonshot AI's Visionary Founder: The Traits Behind His Meteoric Rise

Yang Zhilin, the 34-year-old cofounder of Moonshot AI, has recently captured the attention of the tech world with the la...

Business Insider | Jul 23, 2026, 12:30
Moonshot AI's Visionary Founder: The Traits Behind His Meteoric Rise
View All News