# Research to Reality with Google DeepMind

Paige Bailey, Google DeepMind | AI Engineer World's Fair 2026 | 10:53

Source: https://www.youtube.com/watch?v=zQZiHOpkq_s
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/research-to-reality-with-google-deepmind
Published: 2026-10-05
Tags: edge, open-models, quantization

## TL;DR
- Gemma 4 is an Apache 2-licensed open model family that companies can download, adapt, and fine-tune.
- Gemma can run locally in a browser or on phones, so prompts and data do not need to leave the device.
- Choosing a smaller or quantized model can reduce GPU, laptop, and mobile deployment requirements.

## Summary
Paige Bailey introduces Google DeepMind's work across AI for science, including AlphaFold, robotics, mathematics, and frontier research, then focuses on Gemma 4 and local inference. Gemma 4 comes in sizes from 2 billion to 31 billion parameters and uses an Apache 2 license, allowing companies to download, adapt, and fine-tune the models. Bailey demonstrates Gemma running in a browser through WebAssembly, where it creates an emoji table ranking Harry Potter books without an API call or sending data elsewhere. She then shows Google AI Edge Gallery, which runs downloaded models on Android and iOS for image descriptions, multilingual transcription, function calling, and local skills. The final section connects model size to deployment: larger Gemma variants can run on one commodity GPU, smaller ones can run on laptops or phones, and quantized checkpoints reduce the two-billion-parameter version to less than a gigabyte.

## Key ideas
### DeepMind connects open models with AI for science
[02:38](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=158s)
Bailey frames DeepMind's work around building AI responsibly and for the benefit of humanity. She points to AlphaFold, open models including Gemma, robotics, AI for mathematics, and frontier research. The examples are presented as applications embedded in science projects, rather than as separate demonstrations of model capability. She also invites people working in AI for science to ask how these systems might support their work.

### Managed agents can execute higher-order tasks in a Linux sandbox
[04:01](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=241s)
Bailey describes Google's computer use API and managed agents. A user can describe a higher-order task in natural language, then have a fleet of agents execute it in a Linux workstation-like sandbox. The environment can include added skills and dependencies that are pulled in as needed. She also mentions a speech-to-speech translation API as part of Google's recent product releases.

### Gemma 4 gives companies an open model family to adapt
[04:56](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=296s)
Gemma 4 is presented as Google's latest open model family. Bailey lists 2 billion, 4 billion, 12 billion, 26 billion, and 31 billion parameter versions, including dense and mixture-of-experts models. The models are Apache 2 licensed, so companies can download them, use them in their projects, fine-tune them, and expand on them for their own needs.

### A browser can run Gemma locally with WebAssembly
[05:54](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=354s)
Bailey demonstrates Gemma loading directly in the browser through WebAssembly. The model generates an emoji table comparing Harry Potter books by how funny and exciting they are, then recommends books. The response arrives almost instantly, although Bailey questions its ranking. Because the model is running locally in the browser through transformers.js, the demo does not use an API and the user's data is not sent elsewhere.

### Google AI Edge Gallery puts model skills on phones
[07:35](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=455s)
Google AI Edge Gallery lets users download models and run them on Android and iOS. Bailey lists image description, audio transcription in multiple languages, and on-device function calling. The app also includes skills for building games, writing haikus, querying weather, and scheduling calendar events. Supported device accelerators can perform inference instead of falling back to the CPU, including on higher-end mobile devices such as a Pixel 10.

### Model size determines where local inference can run
[08:59](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=539s)
Bailey compares Gemma's model sizes with deployment needs. She says the 31-billion and 26-billion parameter versions perform above what their sizes might suggest, while also reducing the need for distributed inference and a large GPU footprint. Gemma is intended to run on a single commodity GPU. Some variants run on Jetson Nanos, versions up to 12 billion parameters can run on a laptop, and the 2-billion-parameter model can fit on mobile devices.

### Quantized checkpoints make the smallest models faster and smaller
[10:03](https://www.youtube.com/watch?v=zQZiHOpkq_s&t=603s)
The browser demo used quantization-aware training checkpoints, which Bailey calls QAT checkpoints. These versions reduce the storage required for the model while preserving the local deployment approach. She says the 2-billion-parameter version is less than a gigabyte, helping explain why the browser demo ran so quickly.

## Notable quotes
- "DeepMind's entire mission is to build AI responsibly and for the benefit of humanity." (02:38)
- "Gemma 4 is our open model family." (05:03)
- "This is not using an API. It's just something that's running locally in the browser with transformers.js and the Gemma 4 model." (07:33)
- "You don't have to worry about distributed inference for local models." (09:45)
- "2B can even fit handily on mobile devices." (10:00)

## Tools & references mentioned
- Google DeepMind
- AlphaFold
- Gemma 4
- Gemini 3.5 Flash
- Google AI Edge Gallery
- transformers.js
- WebAssembly
- Apache 2
- Jetson Nano
- Pixel 10

## Who should watch
- You are deciding whether an open model can run inside a company's own application and want concrete model-size and licensing details.
- You need inference to stay in a browser, laptop, or phone instead of sending data to an API.
- You are comparing local deployment options, from a single commodity GPU to mobile hardware, and want to understand the role of quantized checkpoints.

## Related talks

- [Gemma, DeepMind's Family of Open Models](https://aietalks.com/talks/gemma-deepminds-family-of-open-models) (Omar Sanseviero, Google DeepMind, 15:26)
- [Accelerating AI on Edge](https://aietalks.com/talks/accelerating-ai-on-edge) (Chintan Parikh & Weiyi Wang, Google DeepMind, 23:58)
- [Unveiling the Latest Gemma Model Advancements](https://aietalks.com/talks/unveiling-the-latest-gemma-model-advancements) (Kathleen Kenealy, Google DeepMind, 16:25)
- [Gemma 4 Deep Dive](https://aietalks.com/talks/gemma-4-deep-dive) (Cassidy Hardin, Google DeepMind, 19:03)
- [The Desktop Frontier](https://aietalks.com/talks/the-desktop-frontier) (Ahmad Osman, Osmantic, 18:02)
