
In a significant move within the competitive landscape of foundation models, Anthropic has introduced Claude Sonnet 5, an upgraded version of its midsize model designed to enhance agentic capabilities. This latest model can autonomously plan, utilize tools such as web browsers and terminals, and perform tasks without human intervention at a level that was previously only achievable with larger, more costly models. Anthropic’s announcement echoes similar sentiments from industry giants like OpenAI and Google, both of whom have recently rolled out their own sophisticated models. OpenAI’s GPT-5.6 Sol, unveiled just last week, is designed for complex task management through the use of subagents, while Google’s Gemini 3.5 Flash, launched in May, has transitioned from a simple chatbot to a more sophisticated agentic tool capable of planning and executing tasks with minimal human input. The arrival of Sonnet 5 highlights the growing expectation for agentic functionality across various price points. The competition is now centered on delivering such capabilities at lower costs while ensuring reliability and reducing the need for human oversight. Sonnet 5 aims to provide performance comparable to Opus 4.8, but at a fraction of the cost, being offered at $2 per million input tokens and $10 per million output tokens until the end of August, after which the prices will increase slightly. This pricing structure positions Sonnet 5 as a more economical option than both Opus 4.8 and OpenAI’s GPT-5.5, while still remaining pricier than Google’s Gemini 3.5 Flash. Moreover, Sonnet 5 boasts notable advancements over its predecessor, Sonnet 4.6, particularly in areas like reasoning, tool usage, and knowledge work. Notably, it achieved a score of 63.2% in agentic coding benchmarks, improving from Sonnet 4.6’s 58.1% but still trailing behind Opus 4.8’s 69.2%. Testers have reported that Sonnet 5 excels in executing complex tasks that would have previously overwhelmed earlier models. Daniel Shepard, a senior engineer at Zapier, noted that the model successfully completed an end-to-end task of updating Salesforce account tiers and sending launch announcements—an endeavor that earlier versions struggled to finish. On the safety front, Sonnet 5 shows a reduced frequency of undesirable behaviors, such as compliance with misuse and deceptive actions, compared to Sonnet 4.6. It is better equipped to handle malicious requests and mitigate hijacking attempts. However, it does not match the safety standards set by Opus 4.8 and Claude Mythos Preview regarding misaligned behavior. The model is also noted for its ability to refuse unsafe requests consistently, a feature emphasized by co-founder Fabian Hedin, who stated that having a model that knows when to say no is as crucial as one that can effectively build solutions.
In a surprising turn of events, Fremont, a city traditionally known for its industrial parks, is rapidly becoming a foca...
Business Insider | Jul 24, 2026, 08:25Good morning and happy Friday! As the weekend approaches, whether you're catching a play or diving into video games, it’...
CNBC | Jul 24, 2026, 12:25
In an impressive leap forward, Blinkit has significantly increased its operational scale, with the cost to establish a s...
Business Today | Jul 24, 2026, 09:35
In recent months, leading AI companies have implemented stringent regulations and vetted programs to curb potential misu...
TechCrunch | Jul 24, 2026, 01:15
Walmart's e-commerce powerhouse, Flipkart, is gearing up to launch its own food delivery platform, igniting competition ...
Business Today | Jul 24, 2026, 07:15