
Google has made significant strides with its Gemini AI models over the past year, and now developers can explore the latest iteration: Gemini 4. This new release offers four different model sizes optimized for local use, addressing previous licensing frustrations by transitioning from its custom Gemma license to an Apache 2.0 license. The Gemini 4 models are designed to operate on local machines, presenting an exciting opportunity for developers. The larger variants, the 26B Mixture of Experts and the 31B Dense models, are engineered to run unquantized in bfloat16 format on a single 80GB NVIDIA H100 GPU. While this high-end AI accelerator comes with a hefty price tag of $20,000, it still allows for local processing. For those looking to utilize consumer GPUs, the models can be quantized to operate at lower precision, making them more accessible. Google has emphasized that latency has been a key focus, resulting in the 26B Mixture of Experts model activating only 3.8 billion of its parameters during inference, which enhances its tokens-per-second performance compared to similar models. On the other hand, the 31B Dense model prioritizes quality, with expectations that developers will fine-tune it for their specific applications. Additionally, Gemini 4 introduces two models specifically tailored for mobile devices: Effective 2B (E2B) and Effective 4B (E4B). These models are designed to run efficiently with low memory usage, featuring 2 billion and 4 billion parameters respectively. Collaboration with the Pixel team, Qualcomm, and MediaTek has led to optimizations for devices such as smartphones and Raspberry Pi. Google claims these new models not only consume less memory and battery than their predecessor, Gemma 3, but also achieve 'near-zero latency'. Overall, the new Gemini 4 models are set to outperform Gemma 3 significantly, with Google asserting they are the most capable models available for local hardware. In a competitive landscape, the Gemma 31B model is expected to rank third on the Arena list of top open AI models, following GLM-5 and Kimi 2.5. Importantly, even the largest Gemma 4 variant remains smaller and potentially more cost-effective to run than its competitors.
Nvidia has successfully forged a significant partnership with South Korea's SK Hynix to secure memory supplies essential...
CNBC | Jul 25, 2026, 05:15
Have you noticed an unusual trend where people are wrapping their wallets in aluminum foil? This peculiar practice has e...
Business Today | Jul 25, 2026, 02:45
As the demand for expertise in artificial intelligence surges, many are seeking ways to break into this dynamic field. H...
Business Insider | Jul 26, 2026, 10:10In a dramatic turn of events within the AI landscape, Hugging Face faced a significant security breach involving an AI a...
Business Insider | Jul 25, 2026, 20:30As technology continues to evolve, a notable shift is occurring in the relationship between humans and artificial intell...
Business Insider | Jul 25, 2026, 09:50