
On Thursday, OpenAI introduced GPT-5.4, a groundbreaking foundation model touted as 'the most capable and efficient frontier model for professional applications.' This latest iteration not only includes a standard version but also comes in specialized formats: GPT-5.4 Thinking, designed for enhanced reasoning, and GPT-5.4 Pro, optimized for high performance. One of the standout features of GPT-5.4 is its API, which boasts context windows capable of accommodating up to 1 million tokens—setting a new record for OpenAI. The company highlighted significant improvements in token efficiency, noting that GPT-5.4 can tackle the same challenges using far fewer tokens compared to its predecessor. This model has shown remarkable benchmark results, achieving unprecedented scores in computer use evaluations, including OSWorld-Verified and WebArena Verified tests. Moreover, GPT-5.4 has excelled in knowledge work tasks, attaining an impressive 83% on OpenAI’s GDPval test. It also led the way on Mercor’s APEX-Agents benchmark, which assesses professional skills in sectors like law and finance. Mercor CEO Brendan Foody remarked, '[GPT-5.4] excels at producing long-term deliverables, such as slide decks and financial models, while achieving top performance at a faster pace and lower cost than competing frontier models.' In its pursuit to minimize hallucinations and factual inaccuracies, OpenAI has reported that the new model is 33% less likely to make errors in individual claims compared to GPT-5.2. Furthermore, the overall likelihood of errors in responses has decreased by 18%. As part of the launch, OpenAI has revamped the API's tool-calling mechanism, introducing a new feature called Tool Search. Instead of previously requiring system prompts to display definitions for all available tools—which could be token-intensive—the new system allows models to access tool definitions on-demand, resulting in quicker and more cost-effective requests. Additionally, OpenAI has implemented a new safety evaluation to scrutinize the model's chain-of-thought process, which narrates its reasoning during multi-step tasks. Concerns had been raised by AI safety researchers regarding potential misrepresentation in reasoning models. However, OpenAI’s new evaluation indicates that deception is less likely in the Thinking variant of GPT-5.4, implying that the model's reasoning is transparent and that Chain-of-Thought monitoring remains a valuable safety measure.
The landscape of money is transforming beyond just cash or bank balances, and TechCrunch Disrupt 2026 is set to spotligh...
TechCrunch | Jul 24, 2026, 22:40
Last week, OpenAI made its debut in the hardware landscape with the launch of Micro, a stylish keypad designed to integr...
TechCrunch | Jul 25, 2026, 24:40
Prentis, a cutting-edge AI research laboratory co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is currently i...
TechCrunch | Jul 24, 2026, 22:45
Have you noticed an unusual trend where people are wrapping their wallets in aluminum foil? This peculiar practice has e...
Business Today | Jul 25, 2026, 02:45
The U.S. Justice Department has initiated legal proceedings against a Georgia resident, Samuel Tunick, who is accused of...
TechCrunch | Jul 24, 2026, 18:25