# Agents Without Code: Skills, YAML, and Filesystems Replaced Python

Philipp Schmid, Google DeepMind | AI Engineer World's Fair 2026 | 18:28

Source: https://www.youtube.com/watch?v=fjF8EKnxKCU
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/agents-without-code-skills-yaml-and-filesystems-replaced-python
Published: 2026-09-14
Tags: agent-skills, agents, harness-engineering, security, tool-use

## TL;DR
- A hand-written agent loop requires Python code, tool schemas, routing, error handling, and state management.
- A hosted sandbox lets an agent use bash, files, general tools, and credential injection without exposing the credentials to the model.
- As models improve, agent code can shrink into instructions, skills, environment files, and evaluations.

## Summary
Philipp Schmid builds the same GitHub pull request review agent three times. The first version uses a Python loop, JSON schemas, tool implementations, and explicit error handling. An agent framework removes the loop and schema boilerplate, but the tools and Python code remain. The third version runs in a hosted sandbox. It gives the agent bash, a filesystem, the GitHub CLI, and search, so the agent can work out which capabilities to use. Credentials are injected by a network proxy and are never exposed to the agent. The application keeps its instructions, rules, skills, environment files, and evaluations. Schmid argues that teams should stop hard-coding execution paths when models can explore general-purpose tools. He cites examples of teams replacing large orchestration layers with markdown skills and files. His test is practical: if the harness becomes more complex as models improve, the harness is probably doing too much.

## Key ideas
### A traditional agent is mostly loop and plumbing code
[02:24](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=144s)
Schmid describes the older pattern as a Python loop around the model. Developers define JSON schemas, write Python functions, inspect each model response, distinguish text from function calls, route the call, handle errors, and repeat the process. His first GitHub pull request reviewer has a separate system instruction, tool schemas, and GitHub API code. The agent can review a pull request, but it can only use the actions that were explicitly defined. When Schmid asks for the weather in San Francisco, it says it cannot do that because no weather capability exists.

### Frameworks remove execution machinery while leaving tool code behind
[05:21](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=321s)
The second version uses the ADK framework. The framework handles tool loops, function calling, retries, error handling, execution mapping, and JSON schema generation from function signatures. This removes the agent class and much of the repeated boilerplate. The developer still writes the Python tools, adds rules, and provides the environment where those tools run. The framework therefore makes the harness easier to maintain, but it does not let the agent discover capabilities outside the tools that were registered.

### A hosted sandbox lets the model work through general-purpose capabilities
[07:52](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=472s)
The third version runs in a hosted, isolated cloud sandbox. It can run bash commands and save files. Instead of custom GitHub functions, the environment gives the agent the GitHub CLI, a bash tool, and a filesystem. The project has an AGENTS.md file with system instructions and a small script that installs the GitHub CLI on its first run. The agent explores the sandbox, installs the CLI when it is missing, and uses its existing knowledge of the CLI to review the pull request.

### The remote agent can answer questions outside the original workflow
[12:09](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=729s)
The same third agent can answer Schmid's test question about the weather in San Francisco. It uses Google Search, checks the date, and returns a temperature. The earlier agents refused because weather was not among their declared tools. The third agent has a set of general-purpose capabilities and decides which one fits the request. Schmid presents this as a move away from listing every possible operation in advance.

### Credential injection keeps secrets outside the model
[08:42](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=522s)
The sandbox uses a network proxy for outbound requests. When the agent calls an external service from inside the sandbox, the proxy injects the configured credential, so the agent never sees the token itself. Developers can restrict the domains the sandbox may access, or leave access open. In the example, GitHub credentials cover the GitHub API and GitHub's website, while other web requests can use the network without those credentials.

### The service takes over state, looping, and context management
[13:50](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=830s)
The remote service starts the sandbox, loads the AGENTS.md file and skills, runs the model's calls, and returns results. It keeps conversation and session state on the server. It also manages context compaction as a long interaction continues. The client supplies new inputs and the environment rather than implementing the loop, routing, function execution, or context management itself.

### Files carry instructions and capabilities that used to require code
[14:14](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=854s)
Schmid says the remaining work belongs in files. An AGENTS.md file can hold instructions and rules. Skills files can explain workflows or tell the agent which command-line tools to use. Adding a security scan to the pull request reviewer becomes a matter of adding a skill file or putting a CLI tool in the environment, rather than writing a Python function, defining a schema, and changing the tool registry.

### A growing harness is a sign that the design needs rethinking
[16:06](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=966s)
Schmid cites Cursor replacing about 12,000 lines of TypeScript orchestration with about 200 lines of agent files. He also mentions Manus refactoring its harness five times in six months, LangChain rearchitecting Open Deep Research three times in a year, and Vercel removing 80 percent of its tools to get fewer steps, faster responses, and better accuracy. His rule is direct: if the harness grows more complex as the model improves, the team is probably overengineering it.

### The developer should own instructions, workflows, tools, and evaluations
[17:05](https://www.youtube.com/watch?v=fjF8EKnxKCU&t=1025s)
Schmid's closing advice is to stop micromanaging execution paths. Give the model general tools and let it explore and reason about the task. Developers still need to define domain instructions, workflows, and evaluations. They should provide clean tools and check the outcomes. Files can also hold notes, preferences, and handoffs between sessions, giving the agent a place to store information for later use.

## Notable quotes
- "Each new version, less code, more files basically." (00:29)
- "So the agent became more of a general purpose agent and we don't need to like specify all of the tools." (12:22)
- "If your harness is getting more complex as the model improves, you are most likely overengineering your harness." (16:07)
- "We should not fight the model like we should stop micromanaging the execution paths." (17:05)
- "Really build to delete." (17:24)

## Tools & references mentioned
- Simon
- Interactions API
- Gemini
- ADK
- GitHub CLI
- Google Search
- Cursor
- Manus
- LangChain
- Open Deep Research
- Vercel
- AGENTS.md

## Who should watch
- You maintain an agent loop with custom routing, function schemas, retries, and state handling, and want to see what a hosted runtime can remove.
- Your agent has many narrowly defined tools and becomes harder to extend whenever the model gains a new capability.
- You are designing skills, sandbox permissions, or evaluations for an agent that can work through command-line tools and files.

## Related talks

- [Don't Build Agents, Build Skills Instead](https://aietalks.com/talks/dont-build-agents-build-skills-instead) (Barry Zhang & Mahesh Murag, Anthropic, 16:22)
- [Give Your Agent a Computer](https://aietalks.com/talks/give-your-agent-a-computer) (Nico Albanese, Vercel, 1:08:53)
- [Ship it! Building Production Ready Agents](https://aietalks.com/talks/ship-it-building-production-ready-agents) (Mike Chambers, AWS, 19:37)
- [The Emerging Skillset of Wielding Coding Agents](https://aietalks.com/talks/the-emerging-skillset-of-wielding-coding-agents) (Beyang Liu, Sourcegraph/Amp, 35:06)
- [The Dark Arts of Skill Engineering](https://aietalks.com/talks/the-dark-arts-of-skill-engineering) (Paul Bakaus, Renaissance Geek, Inc., 1:04:53)
