# An Interaction Is All You Need

Ivan Leo, Google DeepMind | AI Engineer | 17:04

Source: https://www.youtube.com/watch?v=8aVbXXvJUY4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/an-interaction-is-all-you-need
Published: 2026-10-03
Tags: agents, multimodal, tool-use

## TL;DR
- The Interactions API keeps model context on the server through interaction IDs, so developers do not have to manually carry thought signatures between calls.
- The API uses typed steps and outputs to support multimodal generation, asynchronous tool calls, and combinations of built-in and custom tools.
- Managed Agents provide persistent remote sandboxes with environment IDs, loadable sources, credential-injecting proxies, and named agent configurations.

## Summary
Ivan Leo explains why APIs designed for single model responses become awkward when models reason for longer periods, call several tools, and work across images, video, audio, and external environments. The Interactions API keeps state on the server through interaction IDs and replaces deeply nested response structures with typed outputs and a steps data model. Leo shows how one interaction can move from image generation to video generation while retaining context, and how Google Search, URL Context, and a custom tool can be combined in one call. He then presents Managed Agents, which run the Antigravity harness in persistent remote sandboxes. Environment IDs preserve the sandbox between runs, sources can come from GitHub, GCS, or inline files, and a proxy injects credentials without exposing them to the model. Named agents and the Gemini API CLI extend the same workflow from local development to cloud execution.

## Key ideas
### Agent workloads grew from responses into long-running tool interactions
[00:12](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=12s)
Leo traces a progression from single completions to structured function calls and then to agents that reason, call tools, inspect results, and interact with an environment. Function calling made model output predictable by asking for JSON objects, much like a website sends a registration payload to a backend. Newer models may use several tools and reflect on what each result means before replying. Leo describes an agent as a model with functions, short-term and long-term memory, and an execution environment. A coding agent without somewhere to run code cannot do useful work.

### The Interactions API keeps thought context behind an interaction ID
[04:07](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=247s)
Leo says the older API shape became difficult as Gemini use expanded from ordinary models to agents that can run for minutes and return research reports. Older responses also placed data in deeply nested objects. The Interactions API addresses this with server-side state. On the first turn, the API returns an interaction ID. The client sends that ID as previous interaction ID on the next turn, and the context remains available. This includes the thought signatures used by newer Gemini models. Leo says manually managing those signatures was error-prone, and that losing them can reduce model performance.

### One interaction can carry context across image, video, and audio generation
[06:21](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=381s)
Leo demonstrates a flow that starts with one photograph and creates image variations showing a person in different places. Those calls use interaction IDs, and the same context then feeds video generation with the Omni model. The code passes the returned interaction ID into the next client call and changes the model. Audio and image generation use the same interaction creation method, with the response modality and generation configuration changed for the desired output. Leo says this makes multimodal pipelines easier to build because developers can reuse the same information and context across calls.

### Typed outputs replace difficult nested response parsing
[07:21](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=441s)
The API labels output objects with explicit types. Leo contrasts this with older nested structures that required developers to remember paths such as completions and audio. With a typed output, the client can inspect whether the result is audio or an image and handle it accordingly. Changing the response modality and generation configuration produces an image output without requiring a different general interaction pattern. This structure also supports more complex content as the API handles model responses, tool calls, and generated media in a consistent data model.

### Built-in and custom tools can run together in one interaction
[08:21](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=501s)
Leo shows an agent investigating recent security reports for a React application. The interaction gives the model Google Search, the URL Context tool for retrieving webpage information, and a custom file incident tool. The model decides what information it needs and can call the tools in one API request. This lets the agent work with information beyond its training cutoff because it can access the internet. The example illustrates the API's support for tool combinations rather than separate endpoints or isolated model calls.

### The steps data model describes complex work one step at a time
[09:01](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=541s)
The new steps data model replaces the legacy outputs array with a typed discriminator. Each step can make clear what the model produced, which thought signatures were present, which functions were called, and what kind of content was generated. Leo connects this structure to asynchronous tool calls and models working together. Typed content such as audio and video also fits into the same sequence. The aim is to give developers a data structure that can describe an agent run instead of assuming every request has one message followed by one response.

### Managed Agents provide a persistent Antigravity sandbox
[09:46](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=586s)
Managed Agents run the Antigravity harness in a remote sandbox. Leo says developers previously had to tune a harness, find a sandbox provider, manage infrastructure, and preserve state between runs. With one API call, Managed Agents create a sandbox that can be used repeatedly. In his repository example, the agent lists files, reads them, investigates the code, and produces a report without intervention from the caller. An environment ID routes later requests back to the same sandbox, while an interaction ID preserves the model's conversational context.

### Sources, local skills, and credentials can move with the agent
[11:41](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=701s)
Managed Agents can load Google Cloud Storage buckets, GitHub repositories, and inline files. Installed packages and files created in the sandbox remain available on later turns when the client sends the interaction ID and environment ID. Leo also explains that the same Antigravity agent can run in the local IDE and in the cloud. A developer can package local skills in a folder and upload them as agent sources. For private services, a man-in-the-middle proxy replaces credentials in outbound request headers. The model can use the service without seeing the underlying token.

### Named agents and the Gemini API CLI package local work for reuse
[14:26](https://www.youtube.com/watch?v=8aVbXXvJUY4&t=866s)
Leo describes two ways to create named agents. Developers can freeze configured files and network sources, or chat with an agent while installing dependencies and configuring its environment, then freeze that setup. He says the service supports up to a thousand named agents and charges for the model rather than storage or the sandbox. The open-source Gemini API CLI lets developers test Gemini models locally and package a local coding or Antigravity agent as a folder for upload. An Interactions API skill helps coding agents migrate code and keep model and method information current.

## Notable quotes
- "Fundamentally, if you look at what an agent is, it's really a model that powers a huge chunk of it." (02:02)
- "The interaction ID where you're able to preserve and work with context and ensure you're not busting your cache and the environment ID that enables you to route it to the exact same context." (11:09)
- "The best part about using the Antigravity agent is that fundamentally it's the same agent running in your Antigravity IDE." (12:56)
- "It's only going to execute code that will run against the GitHub API and your token is never exposed." (13:36)

## Tools & references mentioned
- Google DeepMind
- Interactions API
- Managed Agents
- Gemini API
- AI Studio
- Gemini 3.5 Flash
- Gemini 3.1 Flash
- Antigravity
- Antigravity IDE
- Antigravity harness
- Gemini API CLI
- Google Search
- URL Context
- GitHub
- Google Cloud Storage
- Gemini
- Omni model
- Nano Banana
- deep research agent
- React
- Scratch

## Who should watch
- You are building an agent that needs to preserve context across tool calls or across image, video, and audio generation.
- Your current integration manually stores thought signatures or parses deeply nested model responses.
- You need a coding agent with a persistent sandbox, private repository access, reusable skills, or a path from local testing to cloud execution.

## Related talks

- [Building Conversational Agents](https://aietalks.com/talks/building-conversational-agents) (Thor Schaeff & Philipp Schmid, Google DeepMind, 1:47:34)
- [Any-to-Any: Building Native Multimodal Agents](https://aietalks.com/talks/any-to-any-building-native-multimodal-agents) (Patrick Löber, Google DeepMind, 16:21)
- [Milliseconds to Magic: Real-Time Workflows Using the Gemini Live API and Pipecat](https://aietalks.com/talks/milliseconds-to-magic-real-time-workflows-using-the-gemini-live-api-and-pipecat) (Kwindla Kramer, Daily & Shrestha Basu Mallick, Google DeepMind, 21:43)
- [Building Agents (the hard parts!)](https://aietalks.com/talks/building-agents-the-hard-parts) (Rita Kozlov, Cloudflare, 21:12)
- [Useful General Intelligence](https://aietalks.com/talks/useful-general-intelligence) (Danielle Perszyk, Amazon AGI SF Lab, 19:58)
