
In a significant milestone, OpenAI introduced its latest AI model, GPT-5.3-Codex-Spark, on non-Nvidia hardware, specifically designed to run on Cerebras chips. This innovative model boasts an impressive capability of generating code at a staggering pace of over 1,000 tokens per second, marking an approximate 15-fold increase in speed compared to its predecessor. To put this in perspective, Anthropic’s Claude Opus 4.6 in its new premium fast mode achieves around 2.5 times its standard rate of 68.2 tokens per second, although it is a larger and more advanced model than Codex-Spark. "Cerebras has been an exceptional engineering partner, and we are thrilled to introduce rapid inference as a new capability of our platform," stated Sachin Katti, who leads compute operations at OpenAI. The Codex-Spark model is currently available as a research preview to ChatGPT Pro subscribers at $200 per month, accessible through the Codex app, command-line interface, and VS Code extension. OpenAI is also in the process of providing API access to select design partners. Equipped with a 128,000-token context window, the model is designed to process text exclusively at its launch. This latest release builds upon the comprehensive GPT-5.3-Codex model that OpenAI unveiled earlier this month. While the full model is capable of handling more complex coding tasks, Spark is specifically optimized for speed over extensive knowledge. On software engineering benchmarks like SWE-Bench Pro and Terminal-Bench 2.0, Codex-Spark reportedly outperforms the older GPT-5.1-Codex-mini, accomplishing tasks in significantly less time. However, OpenAI has not provided independent validation for these performance claims. Historically, Codex's speed has been a point of contention; a previous test revealed that Codex required nearly twice the time compared to Anthropic’s Claude Code in generating a working version of Minesweeper. In context, GPT-5.3-Codex-Spark's 1,000 tokens per second represents a remarkable advancement over any models OpenAI has previously offered. Independent benchmarks from Artificial Analysis indicate that OpenAI’s fastest models running on Nvidia hardware fall short of this new achievement, with GPT-4 delivering around 147 tokens per second and GPT-4o mini clocking in at approximately 52 tokens per second.
Artificial Intelligence (AI) is transforming our daily experiences, permeating various aspects of technology, including ...
Business Today | Jul 26, 2026, 07:05
This past week has been challenging for the stock market, driven by several significant forces that have created turbule...
CNBC | Jul 25, 2026, 20:05
In the realm of cybersecurity, few figures are as intriguing as Phineas Fisher, a hacker who has evaded capture for near...
TechCrunch | Jul 25, 2026, 21:00
In recent discussions, a once-obscure topic in artificial intelligence has surged to the forefront of debates among tech...
CNBC | Jul 25, 2026, 12:15
As technology continues to evolve, a notable shift is occurring in the relationship between humans and artificial intell...
Business Insider | Jul 25, 2026, 09:50