Agents Without Code: Skills, YAML, and Filesystems Replaced Python

Philipp Schmid, Google DeepMind18:28 · Sept 2026 · 15K views
Thumbnail for Agents Without Code: Skills, YAML, and Filesystems Replaced Python Watch on YouTube
TL;DR
  1. 1

    A hand-written agent loop requires Python code, tool schemas, routing, error handling, and state management.

  2. 2

    A hosted sandbox lets an agent use bash, files, general tools, and credential injection without exposing the credentials to the model.

  3. 3

    As models improve, agent code can shrink into instructions, skills, environment files, and evaluations.

Summary

Philipp Schmid builds the same GitHub pull request review agent three times. The first version uses a Python loop, JSON schemas, tool implementations, and explicit error handling. An agent framework removes the loop and schema boilerplate, but the tools and Python code remain. The third version runs in a hosted sandbox. It gives the agent bash, a filesystem, the GitHub CLI, and search, so the agent can work out which capabilities to use. Credentials are injected by a network proxy and are never exposed to the agent. The application keeps its instructions, rules, skills, environment files, and evaluations. Schmid argues that teams should stop hard-coding execution paths when models can explore general-purpose tools. He cites examples of teams replacing large orchestration layers with markdown skills and files. His test is practical: if the harness becomes more complex as models improve, the harness is probably doing too much.

Key ideas
02:24

A traditional agent is mostly loop and plumbing code

Schmid describes the older pattern as a Python loop around the model. Developers define JSON schemas, write Python functions, inspect each model response, distinguish text from function calls, route the call, handle errors, and repeat the process. His first GitHub pull request reviewer has a separate system instruction, tool schemas, and GitHub API code. The agent can review a pull request, but it can only use the actions that were explicitly defined. When Schmid asks for the weather in San Francisco, it says it cannot do that because no weather capability exists.

05:21

Frameworks remove execution machinery while leaving tool code behind

The second version uses the ADK framework. The framework handles tool loops, function calling, retries, error handling, execution mapping, and JSON schema generation from function signatures. This removes the agent class and much of the repeated boilerplate. The developer still writes the Python tools, adds rules, and provides the environment where those tools run. The framework therefore makes the harness easier to maintain, but it does not let the agent discover capabilities outside the tools that were registered.

07:52

A hosted sandbox lets the model work through general-purpose capabilities

The third version runs in a hosted, isolated cloud sandbox. It can run bash commands and save files. Instead of custom GitHub functions, the environment gives the agent the GitHub CLI, a bash tool, and a filesystem. The project has an AGENTS.md file with system instructions and a small script that installs the GitHub CLI on its first run. The agent explores the sandbox, installs the CLI when it is missing, and uses its existing knowledge of the CLI to review the pull request.

12:09

The remote agent can answer questions outside the original workflow

The same third agent can answer Schmid's test question about the weather in San Francisco. It uses Google Search, checks the date, and returns a temperature. The earlier agents refused because weather was not among their declared tools. The third agent has a set of general-purpose capabilities and decides which one fits the request. Schmid presents this as a move away from listing every possible operation in advance.

08:42

Credential injection keeps secrets outside the model

The sandbox uses a network proxy for outbound requests. When the agent calls an external service from inside the sandbox, the proxy injects the configured credential, so the agent never sees the token itself. Developers can restrict the domains the sandbox may access, or leave access open. In the example, GitHub credentials cover the GitHub API and GitHub's website, while other web requests can use the network without those credentials.

13:50

The service takes over state, looping, and context management

The remote service starts the sandbox, loads the AGENTS.md file and skills, runs the model's calls, and returns results. It keeps conversation and session state on the server. It also manages context compaction as a long interaction continues. The client supplies new inputs and the environment rather than implementing the loop, routing, function execution, or context management itself.

14:14

Files carry instructions and capabilities that used to require code

Schmid says the remaining work belongs in files. An AGENTS.md file can hold instructions and rules. Skills files can explain workflows or tell the agent which command-line tools to use. Adding a security scan to the pull request reviewer becomes a matter of adding a skill file or putting a CLI tool in the environment, rather than writing a Python function, defining a schema, and changing the tool registry.

16:06

A growing harness is a sign that the design needs rethinking

Schmid cites Cursor replacing about 12,000 lines of TypeScript orchestration with about 200 lines of agent files. He also mentions Manus refactoring its harness five times in six months, LangChain rearchitecting Open Deep Research three times in a year, and Vercel removing 80 percent of its tools to get fewer steps, faster responses, and better accuracy. His rule is direct: if the harness grows more complex as the model improves, the team is probably overengineering it.

17:05

The developer should own instructions, workflows, tools, and evaluations

Schmid's closing advice is to stop micromanaging execution paths. Give the model general tools and let it explore and reason about the task. Developers still need to define domain instructions, workflows, and evaluations. They should provide clean tools and check the outcomes. Files can also hold notes, preferences, and handoffs between sessions, giving the agent a place to store information for later use.

"If your harness is getting more complex as the model improves, you are most likely overengineering your harness."16:07
Who should watch
  • You maintain an agent loop with custom routing, function schemas, retries, and state handling, and want to see what a hosted runtime can remove.
  • Your agent has many narrowly defined tools and becomes harder to extend whenever the model gains a new capability.
  • You are designing skills, sandbox permissions, or evaluations for an agent that can work through command-line tools and files.