Gemma, DeepMind's Family of Open Models

Omar Sanseviero, Google DeepMind15:26 · Apr 2026 · 26K views
Thumbnail for Gemma, DeepMind's Family of Open Models Watch on YouTube
TL;DR
  1. 1

    Gemma 4 ranges from models that run on phones and Raspberry Pis to a 31-billion-parameter model that fits on a consumer GPU.

  2. 2

    Per-layer embeddings let the E2B and E4B models keep only part of their parameters in the GPU, which suits mobile devices.

  3. 3

    Gemma's open ecosystem includes multilingual fine-tunes, on-device applications, safety and medical variants, and research in areas such as sovereign AI and cancer therapy.

Summary

Omar Sanseviero introduces Gemma 4, Google DeepMind's latest family of open models. The range goes from 2 billion to 32 billion parameters, with small multimodal reasoning models for phones and Raspberry Pis, a mixture-of-experts model for low latency, and a 31-billion-parameter model for higher capability on consumer GPUs. He demonstrates offline coding, agentic tasks, and parallel SVG generation on local devices. He explains the E2B architecture, where per-layer embeddings reduce the amount of data that must stay in GPU memory. Gemma 4 supports images, video, audio, speech translation, and more than 140 languages. Sanseviero also covers the Apache 2 license, integrations with Android Studio and llama.cpp, community fine-tunes, Shield Gemma, MedGemma, and sovereign-language projects. He argues that open models are becoming practical for private, offline, customized applications, while acknowledging that API models remain preferable when maximum raw intelligence is the priority.

Key ideas
01:07

Gemma 4 covers phones, consumer GPUs, and larger local deployments

Gemma 4 ranges from 2 billion to 32 billion parameters. The smallest two models can run on Android phones, iPhones, and Raspberry Pis, while supporting multimodality, reasoning, and on-device agentic tasks. A mixture-of-experts model targets high speed and low latency. The 31B model is the most capable option in the family, yet Sanseviero says it can still run on a consumer GPU. The models were deliberately kept in developer-friendly sizes, so users can choose between local device deployment and greater raw capability without moving to a large server-only model.

02:25

Local demonstrations show coding and agents running without network access

Sanseviero shows Gemma running directly on an Android phone, where an agent can select skills such as playing the piano. A coding example also runs on the phone in airplane mode, with no API calls. On a laptop, 10 Gemma instances run in parallel through llama.cpp, each generating a different SVG. The results appear within seconds, at about 100 tokens per second. He presents these examples as evidence that Gemma can handle coding, Android app development, and agentic tasks entirely on local hardware.

03:47

Gemma 4 improves capability without requiring larger machines

Sanseviero uses LM Arena scores as a rough proxy for how much the community likes models in general conversation and related tasks, while noting that it is not a perfect benchmark. Gemma models appear in the upper-left area of his chart, which indicates small models with strong scores. He says Gemma 1, Gemma 2, Gemma 3, and Gemma 4 have continued to improve without simply becoming bigger. He expects increasingly capable models to run directly on personal devices, including phones.

05:02

The Apache 2 license gives users more freedom to use Gemma 4

Users had told Google that the earlier Gemma licenses were not sufficiently open. With Gemma 4, Google changed the license to Apache 2. Sanseviero describes this as giving users the flexibility and control associated with that license. The change matters for teams that want to download the models, run them in their own infrastructure or devices, and adapt them for their own applications.

05:27

Per-layer embeddings reduce GPU memory pressure on mobile devices

E2B means effectively 2 billion parameters, although the model has about 4 billion parameters in total. Its per-layer embeddings use a lookup-table-like structure rather than the matrix computations associated with ordinary transformer layers. The embeddings can sit in CPU memory or on disk, while the GPU loads only about 2 billion parameters. Sanseviero says this architecture is designed for mobile use cases. With llama.cpp, users can move the per-layer embeddings to the CPU or disk using an override-tensor flag.

06:57

Gemma 4 combines multimodal input with broad language coverage

The smallest models understand images, video, and audio. They can perform speech recognition and translate spoken Spanish into text in another language. Larger models can locate objects such as a llama in an image, detect multiple objects, and explain images containing Japanese text. Gemma 4 was trained with more than 140 languages and uses a tokenizer based on Gemini. Sanseviero says the tokenizer was designed for multilingual use, which can help when fine-tuning for lower-resource languages such as Quechua.

08:42

The Gemma ecosystem grows through community tools and fine-tunes

One week after release, Gemma 4 had reached 10 million downloads for its base models and had more than 1,000 community models, including quantizations and fine-tunes. The wider Gemma family had passed 500 million downloads. Sanseviero names Unsloth, MLX, llama.cpp, Hugging Face, vLLM, and C Lang as ecosystem collaborators. Google wants new Gemma tools to work through the tools developers already use, such as Hugging Face Transformers, rather than requiring a switch to Keras.

11:27

Official and community variants target safety, medicine, and sovereign language needs

Shield Gemma is a family of models for detecting toxic images or text against application policies. MedGemma is a multimodal Gemma 3-based model for medical tasks including radiology and chest X-ray understanding, and users can fine-tune it for narrower needs. Sanseviero also describes AI Singapore's work on Southeast Asian languages and Sarvam's work in India, where national-language model efforts connect to sovereign AI goals. He cites DeepMind research that used Gemma 3 to propose cancer therapy pathways later validated in a lab.

"Open models means that these are models that you can take, you can download, you can run in your own infrastructure, your own devices."00:07
Who should watch
  • You are deciding whether a small open model can run on a phone, laptop, Raspberry Pi, or other local hardware.
  • Your application needs offline inference, private data handling, multilingual input, or local customization.
  • You are evaluating Gemma fine-tuning, Android development, safety filtering, medical models, or sovereign-language projects.