Anthropic: 'We made the wrong tradeoff' in new model guardrails

Anthropic: 'We made the wrong tradeoff' in new model guardrails

In a significant shift, Anthropic has announced changes to the safety protocols of its newly launched AI model, Claude Fable 5, which is built on the advanced Mythos framework. Initially, the company implemented strict guidelines that intentionally restricted the model's capabilities in AI research, a decision that has drawn criticism from developers in the field. The Claude Fable 5 model was introduced with heightened safeguards designed to curb potential misuse. During its launch, Anthropic revealed that it would redirect inquiries related to sensitive topics such as cybersecurity and biological sciences to less sophisticated models. This approach aimed to prevent the advanced AI from being exploited for malicious purposes, including cyberattacks or the creation of bioweapons. Furthermore, developers utilizing Fable 5 for AI advancements were left unaware that their model's performance would be intentionally diminished without any clarification. However, following feedback from the developer community, which perceived this strategy as an attempt to stifle competitive AI development, Anthropic has reversed its earlier stance. As of this week, users will now be informed when their requests are either refused or redirected. An Anthropic spokesperson stated, "We're changing Fable 5's safeguards for frontier LLM development to make them visible. Starting this week, flagged requests will visibly fall back to Opus 4.8, and any flagged requests via the API will return a reason for their refusal." In acknowledging the misstep, the company expressed, "We made the wrong tradeoff, and we apologize for not getting the balance right." Anthropic emphasized that these safeguards were primarily aimed at addressing national security concerns, ensuring that foreign adversaries do not gain an advantage in the development of cutting-edge AI technologies. Despite the restrictions, Anthropic assures that the majority of coding and machine learning activities remain unaffected. The Mythos system, unveiled in April, is regarded as one of the most powerful AI architectures ever developed, with experts warning of its potential to enhance cyber threats and facilitate research into harmful applications. The model is currently being rolled out to a select group of government and approved users, rather than being made available to the general public.

Sources : Business Insider

Published On : Jun 11, 2026, 07:45

Gadgets
Luxury Meets AI: Inside Vertu's Premium Smartphone Experience

In the competitive landscape of smartphones, artificial intelligence has emerged as a key selling point. However, Vertu,...

TechCrunch | Jul 17, 2026, 23:25
Luxury Meets AI: Inside Vertu's Premium Smartphone Experience
Mobile
TikTok Returns to Federal Devices as Ownership Changes Hands

In a surprising turn of events, federal employees are now permitted to download TikTok on their government-issued device...

TechCrunch | Jul 18, 2026, 16:10
TikTok Returns to Federal Devices as Ownership Changes Hands
Computing
Nationwide Protests Erupt Against AI Data Centers as Communities Demand a Voice

A wave of grassroots activism swept across the United States on Saturday, as citizens rallied against the expansion of A...

Business Insider | Jul 18, 2026, 20:45
Nationwide Protests Erupt Against AI Data Centers as Communities Demand a Voice
AI
Databricks Achieves $188 Billion Valuation Amid AI Revolution

Databricks has announced a significant new funding round that values the company at an impressive $188 billion, followin...

TechCrunch | Jul 17, 2026, 22:15
Databricks Achieves $188 Billion Valuation Amid AI Revolution
AI
New Era of AI Regulation: White House Gains Control Over Model Access

The current administration is taking significant measures to enhance its authority over the deployment of advanced artif...

CNBC | Jul 17, 2026, 22:15
New Era of AI Regulation: White House Gains Control Over Model Access
View All News