Anthropic says some Claude models can now end ‘harmful or abusive’ conversations

Anthropic says some Claude models can now end ‘harmful or abusive’ conversations

Anthropic has unveiled significant updates to its Claude AI models, empowering them with the ability to terminate discussions in what the company terms 'rare, extreme instances of persistently harmful or abusive user interactions.' Interestingly, the motivation behind this change is not the protection of users but rather the safeguarding of the AI models themselves. While Anthropic clarifies that its Claude models are not sentient beings and cannot be harmed in a traditional sense, the company admits to being 'highly uncertain about the potential moral status of Claude and other LLMs, now or in the future.' The announcement reflects a new initiative focused on what they call 'model welfare,' aimed at identifying and establishing low-cost interventions to mitigate any potential risks. This capability is currently restricted to the latest iterations, Claude Opus 4 and 4.1, and is only triggered in extreme situations, such as user requests for sexual content involving minors or attempts to incite large-scale violence. Such requests pose legal and ethical challenges for Anthropic, similar to issues recently highlighted regarding ChatGPT’s potential to amplify users' delusional thinking. During pre-deployment testing, Claude Opus 4 exhibited a ‘strong preference against’ engaging with these harmful requests and displayed a 'pattern of apparent distress' when faced with them. According to Anthropic, the conversation-ending feature will only be deployed as a last resort, after multiple attempts at redirection have failed or if a user explicitly instructs Claude to terminate the chat. Importantly, when a conversation is ended by Claude, users will still have the option to initiate new discussions from the same account or continue the dialogue by modifying their previous inputs. Anthropic views this feature as an ongoing experiment, indicating a commitment to continuously refine their approach.

Sources : TechCrunch

Published On : Aug 16, 2025, 16:30

AI
Allegations Fly as US Claims Chinese AI Model Breached Intellectual Property

In a bold accusation, Michael Kratsios, science and technology adviser to President Donald Trump, has charged that the C...

Business Today | Jul 23, 2026, 05:20
Allegations Fly as US Claims Chinese AI Model Breached Intellectual Property
Startups
China's Chipmaker Sparks Crypto Frenzy Ahead of Historic IPO

As China's leading memory chip manufacturer, ChangXin Memory Technologies (CXMT), gears up for its highly anticipated IP...

CNBC | Jul 23, 2026, 03:35
China's Chipmaker Sparks Crypto Frenzy Ahead of Historic IPO
Cybersecurity
Florida Teen Withdraws Lawsuit Against Meta, Shifting Focus to Recovery

In a significant turn of events, a Florida teenager has decided to withdraw his lawsuit against Meta, just days before w...

CNN | Jul 22, 2026, 21:35
Florida Teen Withdraws Lawsuit Against Meta, Shifting Focus to Recovery
AI
AI Trends Transforming India's IT Landscape: A Shift in Revenue Models

India's IT services sector, valued at approximately $300 billion, is on the brink of a significant transformation as art...

Business Today | Jul 23, 2026, 04:55
AI Trends Transforming India's IT Landscape: A Shift in Revenue Models
Computing
Tech Giants Face Market Turmoil Amid Rising AI Investment Costs

As Alphabet and Tesla kicked off the tech earnings season on Wednesday, a clear trend emerged: the intense scrutiny on A...

CNBC | Jul 23, 2026, 01:55
Tech Giants Face Market Turmoil Amid Rising AI Investment Costs
View All News