Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster

Google’s Gemma 4 open AI models use “speculative decoding” to get up to 3x faster

This spring, Google unveiled its Gemma 4 open models, setting a new standard for local AI performance. Now, the tech giant is enhancing these capabilities further with the introduction of Multi-Token Prediction (MTP) drafters for the Gemma series. Google's MTP employs a novel technique known as speculative decoding, allowing the models to predict future tokens during generation. This innovation significantly accelerates the token generation process compared to traditional methods where tokens are produced in a linear sequence. The latest iteration of Gemma retains the foundational technology that drives Google’s advanced Gemini AI but is specifically optimized for local deployment. This means that users can run these models on their hardware without needing to rely on cloud-based systems, thus keeping their data within their control. Gemini itself is designed to work seamlessly with Google’s custom TPU chips, which are capable of operating in large clusters with high-speed interconnects and memory. With a single high-performance AI accelerator, users can run the full precision of the largest Gemma 4 model, while quantization techniques enable it to operate on standard consumer GPUs. While the Gemma models offer exciting opportunities for local AI experimentation, Google has also recognized the limitations of typical consumer hardware. This is where MTP becomes vital. Traditional large language models (LLMs) like Gemma generate tokens one at a time, leading to inefficiencies, especially when processing power is underutilized. MTP optimizes this process by using lighter draft models, which, despite having only 74 million parameters in the Gemma 4 E2B, are fine-tuned to facilitate faster speculative token generation. By sharing the key value cache—essentially the model's active memory—these drafters avoid the need for repetitive context recalculation, thus streamlining the generation process. Additionally, the E2B and E4B drafters utilize a sparse decoding method to efficiently identify clusters of probable tokens, enhancing overall performance.

Sources : Ars Technica

Published On : May 06, 2026, 15:50

Science
Vision Breakthrough: US Startup Launches Groundbreaking Retina Chip in Europe

Science Corp is poised to introduce a revolutionary retina chip in Europe, designed to restore partial vision for indivi...

Business Today | Jul 25, 2026, 01:00
Vision Breakthrough: US Startup Launches Groundbreaking Retina Chip in Europe
AI
Hugging Face CEO Calls for Action Following AI Security Breach

In a dramatic turn of events within the AI landscape, Hugging Face faced a significant security breach involving an AI a...

Business Insider | Jul 25, 2026, 20:30
Hugging Face CEO Calls for Action Following AI Security Breach
AI
Revolutionizing Office Automation: Prentis Aims to Secure $100 Million in Funding

Prentis, a newly established AI research lab, is making waves in the tech industry as it prepares to raise $100 million ...

TechCrunch | Jul 25, 2026, 24:00
Revolutionizing Office Automation: Prentis Aims to Secure $100 Million in Funding
Computing
Market Turbulence: Four Key Factors Impacting Stocks This Week

This past week has been challenging for the stock market, driven by several significant forces that have created turbule...

CNBC | Jul 25, 2026, 20:05
Market Turbulence: Four Key Factors Impacting Stocks This Week
Cybersecurity
The Elusive Phineas Fisher: The Hacktivist Who Took Down Spyware Giants

In the realm of cybersecurity, few figures are as intriguing as Phineas Fisher, a hacker who has evaded capture for near...

TechCrunch | Jul 25, 2026, 21:00
The Elusive Phineas Fisher: The Hacktivist Who Took Down Spyware Giants
View All News