
Google has introduced Gemma 4, the latest iteration of its open AI model family, designed to enhance reasoning and agent-like capabilities across various devices, from smartphones to powerful enterprise GPUs. Gemma 4 represents a significant advancement in Google's open model offerings, developed using the same foundational research as the proprietary Gemini models. Unlike traditional, large-scale closed systems, Gemma aims to provide a lightweight and versatile solution that can be deployed on a broad spectrum of hardware configurations. This initiative seeks to bridge the gap between open and proprietary AI ecosystems, offering developers the flexibility to build applications locally or leverage cloud infrastructure for scaling. The Gemma 4 family introduces four different model variants, with Google asserting that the larger models rank among the leading open models worldwide. Remarkably, these models outperform much larger systems while utilizing fewer parameters. The 26B MoE model enhances speed and efficiency by activating only a portion of its parameters during inference, whereas the 31B Dense model prioritizes output quality to ensure maximum performance. A notable feature of Gemma 4 is its emphasis on 'intelligence-per-parameter.' Rather than simply increasing model size, these new models have been optimized to deliver superior performance with reduced computational demands. The smaller E2B and E4B models are specifically designed for on-device applications, catering to smartphones, IoT devices, and embedded systems. These models support various input types, including audio, images, and video, and are engineered to operate offline with minimal latency. Gemma 4 also brings several enhancements tailored for developers and enterprises. Released under the permissive Apache 2.0 license, it allows for commercial use, modification, and deployment without stringent limitations. The models are readily available on platforms like Hugging Face, Kaggle, and Ollama, and they are compatible with a variety of tools, including Transformers, vLLM, and llama.cpp. Developers can easily fine-tune these models locally or deploy them through cloud services, maximizing their application potential.
On Friday evening, SpaceX executed a significant milestone by launching its colossal Starship rocket from its facility i...
CNBC | Jul 25, 2026, 24:10
Prentis, a newly established AI research lab, is making waves in the tech industry as it prepares to raise $100 million ...
TechCrunch | Jul 25, 2026, 24:00
The landscape of money is transforming beyond just cash or bank balances, and TechCrunch Disrupt 2026 is set to spotligh...
TechCrunch | Jul 24, 2026, 22:40
Investors are growing increasingly uneasy about the substantial capital required to realize the ambitions of artificial ...
CNBC | Jul 24, 2026, 20:15
Have you noticed an unusual trend where people are wrapping their wallets in aluminum foil? This peculiar practice has e...
Business Today | Jul 25, 2026, 02:45