Google’s latest DiffusionGemma open AI model comes with a 4x speed boost

Google’s latest DiffusionGemma open AI model comes with a 4x speed boost

In a notable advancement in artificial intelligence, Google DeepMind has introduced DiffusionGemma, a groundbreaking addition to the Gemma 4 open model family. This innovative model departs from traditional linear text generation methods, enabling it to produce entire blocks of text simultaneously. The key to this enhanced speed and efficiency lies in DiffusionGemma's unique architecture, which allows it to operate effectively on local hardware, including Nvidia DGX systems and standard gaming GPUs. Unlike most AI models that generate text in a sequential, left-to-right manner, DiffusionGemma employs a technique reminiscent of image generation models. It begins with a field of placeholder tokens and iteratively refines them to produce the final output, resulting in what is termed a “denoised” text canvas. With 26 billion parameters, DiffusionGemma is classified as a Mixture of Experts (MoE) model. However, only 3.8 billion parameters are activated during inference, ensuring it can run efficiently within the 18GB RAM limits of high-end GPUs. In testing scenarios using an RTX 5090, the model achieves an impressive output of approximately 700 tokens per second. When paired with a single Nvidia H100 AI accelerator, it can exceed 1,000 tokens per second—about four times the output of its autoregressive counterparts. This innovative approach shifts the focus from memory bandwidth limitations to computational power, allowing for the generation of up to 256 tokens in parallel. Google highlights that this capability significantly enhances performance in complex tasks such as in-line editing, molecular sequencing, and mathematical graphing. A demonstration of DiffusionGemma's prowess is evident in its ability to solve Sudoku puzzles—an area where standard autoregressive models often struggle due to the interdependencies of tokens. By continuously self-correcting large sets of tokens, DiffusionGemma streamlines this challenging process.

Sources : Ars Technica

Published On : Jun 10, 2026, 19:30

Space
SpaceX Successfully Tests Starship Rocket, Launching New Era of Space Exploration

On Friday evening, SpaceX executed a significant milestone by launching its colossal Starship rocket from its facility i...

CNBC | Jul 25, 2026, 24:10
SpaceX Successfully Tests Starship Rocket, Launching New Era of Space Exploration
Aerospace
SpaceX's Latest Starlink Launch Marks Milestone Amid Booster Setbacks

On Friday, SpaceX marked a significant achievement by successfully launching its first batch of third-generation Starlin...

TechCrunch | Jul 24, 2026, 23:40
SpaceX's Latest Starlink Launch Marks Milestone Amid Booster Setbacks
Computing
Crisis Averted: Power Line Failure Highlights Urgent Need for Data Center Resilience

A power line failure near Washington, DC, recently showcased a significant challenge faced by the electrical grid due to...

TechCrunch | Jul 25, 2026, 13:50
Crisis Averted: Power Line Failure Highlights Urgent Need for Data Center Resilience
AI
Prentis: Ambitious New AI Lab Eyes $100 Million Funding with Innovative Automation Goals

Prentis, a cutting-edge AI research laboratory co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is currently i...

TechCrunch | Jul 24, 2026, 22:45
Prentis: Ambitious New AI Lab Eyes $100 Million Funding with Innovative Automation Goals
Cybersecurity
Vietnam's Controversial Social Media Proposal: A New Approach for Youth Engagement

Vietnam is contemplating a distinctive approach to youth social media regulations, diverging from the more common outrig...

TechCrunch | Jul 24, 2026, 21:25
Vietnam's Controversial Social Media Proposal: A New Approach for Youth Engagement
View All News