Google DeepMind officially launched Gemma 4 on April 2, 2026, describing it as "byte by byte, the most performant open models ever released." Built from Gemini 3 research and technology, this new model family is available under an Apache 2.0 license, meaning any developer, researcher, or company can freely use, modify, and deploy them. Gemma 4 comes in four sizes: 2B and 4B for mobile and IoT devices, and 26B and 31B for more demanding workloads on personal hardware or servers. Clement Farabet, VP Research at Google DeepMind, and Olivier Lacombe, Group Product Manager, jointly presented the model family.

Gemma 4's performance represents a generational leap over Gemma 3. The 31B model achieves an 89.2% score on the AIME benchmark (a demanding mathematical reasoning test), compared to approximately 60% for its predecessor – a "genuinely shocking" gain according to independent analysts. The models natively integrate agentic capabilities: function calling, tool usage, and multi-step complex task orchestration. According to evaluators, Gemma 4 26B surpasses much larger proprietary models in numerous coding, reasoning, and multilingual comprehension tests.

Gemma 4's architecture is hailed by researchers as "the most interesting of the year" among open models. Google DeepMind has maximized intelligence per parameter – a crucial metric for deployment in edge computing, on smartphones, or in resource-constrained environments. The 2B and 4B models offer unprecedented intelligence levels for such compact sizes, paving the way for truly autonomous embedded AI applications. Compatibility with tools like Ollama, llama.cpp, and vLLM facilitates local deployment.

In the competitive open-source AI landscape, Gemma 4 directly challenges Meta's Llama 4 and Alibaba's Qwen 3.6. Google's Apache 2.0 license is considered the most permissive on the market, whereas Meta imposes restrictions on Llama for applications exceeding 700 million users. Google is betting on the ecosystem: Gemma 4 is accessible via Google AI Studio, Hugging Face, Kaggle, and Vertex AI, with deployment tutorials on each platform. Google's strategy is clear: by making its models open and performant, it draws developers into its cloud ecosystem while accelerating global AI adoption.

The launch of Gemma 4 occurs during an exceptional week for AI: Anthropic unveiled Claude Mythos, OpenAI closed its $122 billion funding round, and OpenAI published its policy blueprint "Industrial Policy for the Intelligence Age" calling for robot taxes, AI sovereign wealth funds, and a four-day work week. Open-source AI has never been more performant, but the question of monetization remains open: how does Google profit from models it distributes for free? The answer, as always with Google, lies in the cloud – and in the conviction that whoever controls the models controls the infrastructure. Sources and links: - Google Blog: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/ - Google DeepMind: https://deepmind.google/models/gemma/gemma-4/ - Towards AI: https://pub.towardsai.net/googles-gemma-4-is-the-most-architecturally-interesting-open-model-released-this-year-b245a406cd6a - AI Workflows: https://www.aiworkflows.tools/blog/google-gemma-4-complete-guide-open-weight-agentic-ai-2026 - Medium (Benchmarks): https://medium.com/@moksh.9/heres-a-tighter-benchmark-focused-blog-post-501c5ea829f4