An agent with no prewritten task-specific tools can create a calculator or character-counter tool at runtime and use it without restarting.
2
Strands Agents enables runtime tool creation with an editor, shell, and load_tool, guided by a system prompt that defines the tool format and rules.
3
Self-modifying agents need evals for goal success, tool use, parameters, inter-agent flow, and guardrails around execution, permissions, and observability.
Summary
Sandhya Subramani demonstrates an agent that begins with no task-specific tools and writes the tools it needs while running. In her examples, the agent creates a math calculator and a character counter, loads each new tool, and uses it without a restart. She explains how Strands Agents supports this pattern with an editor, shell, and load_tool, plus a system prompt that describes valid tools, storage rules, and when to create them. The same approach can create sub-agents for a travel planner, using patterns such as swarm, graph, handoff, and workflow. Subramani is direct about the risks. A system that can create tools and agents can also modify or delete things. She recommends evaluating the end goal, answers, tool choices, parameters, call sequence, and inter-agent messages. Sandboxed execution, restricted permissions, and observability are required before using this pattern in production.
Runtime tool creation lets an agent recover from missing capabilities
Subramani contrasts ordinary coding workflows with an agent that can repair itself while running. In her example, an agent encounters a request for a complex calculation even though its tools directory contains no task-specific tools. It writes a math calculator, creates the file during the active session, and uses it without rerunning the program. When asked to count characters, it notices that its existing tools cannot do that, writes a character counter, loads it, and applies it to the input. The point is that the agent can respond to a missing capability instead of stopping for a developer to add code and restart the service.
Strands Agents needs three runtime tools and a tool-writing prompt
Strands Agents is the open-source agentic harness Subramani uses for the demo. The runtime setup gives the agent an editor, a shell, and load_tool. The editor lets it write files, the shell gives it information about the directory and execution context, and load_tool loads generated tools dynamically. A system prompt defines the tool template, explains where tools should be stored, and gives rules such as checking whether a tool already exists before writing a new one. The agent can use different model providers without changing the surrounding architecture.
The system prompt turns tool generation into a repeatable operation
The prompt does more than tell the model to write code. It says that the agent is a meta-tooling agent, describes what a valid tool looks like, specifies the tool format, and gives the location where generated code belongs. Subramani also tells it to check for an existing tool first and create a new one only when needed. With this prompt and the three runtime tools, a request such as creating five random tools becomes an executable operation. The generated files can then be loaded and used during the same conversation.
The same mechanism can assemble specialist sub-agents
Subramani says an agent can create agents as well as tools. She describes swarm patterns, where sub-agents work on parts of a task in parallel, graph patterns, where one result feeds another agent, handoffs, and workflows that combine these structures. Her travel-planner demo uses a prompt that breaks the task into two to four sub-agents. For a Hawaii itinerary, the agent creates a flight agent, an activities agent, and an itinerary agent, then calls them. The demo does not have real-time information because no tool with access to a real-world API was provided.
Self-healing requires the agent to inspect and update its own failures
The generated agents can also be asked to repair themselves. Subramani describes deliberately pushing the demo into errors, where the agent recognizes that it is failing and edits the same tool or agent to fix the problem. She also describes changing the request after the first plan, such as replacing Hawaii with another destination, and having the agent update its work. The capability is useful because the system can respond to failures and changed requirements during execution. It also means the agent is making code changes autonomously, so the failure modes need to be measured and controlled.
Evals must cover outcomes, tool calls, parameters, and agent communication
Subramani describes eight built-in evals in Strands Agents. At the session level, one question is whether the agent achieved the requested outcome, such as booking a flight. At the trace level, evaluators inspect whether the answer was useful and whether it was made up. They also check whether the agent selected the right tool and passed the right parameter, such as the correct account ID. In multi-agent systems, evaluation includes the final result, the order of tool and sub-agent calls, and the messages passed between agents.
Sandboxing and restricted permissions limit the damage of self-modification
A self-modifying agent can create tools and agents, but it can also modify or delete files. Subramani recommends sandboxing the environment where generated code executes, rather than protecting only the agent container from the outside world. She also calls for constrained tool permissions, limits on who can ask the agent to perform actions, and limits on the types of requests it accepts. Observability matters as well. Evals, telemetry, and related execution data give the team evidence about what the system did before granting it broader autonomy.
Strands has moved from generated tools to generated agents and source updates
Subramani presents runtime tool creation as an early stage of self-improving agents. Strands first supported agents that built their own tools, then agents that created other agents, and now agents that can update source code. She says AWS originally built Strands internally and later open-sourced it. The Python version then wrote the TypeScript version of Strands by itself, which she gives as an example of the system updating its own source code.
"But imagine if your agent was smart enough to figure out, oh man, there is a bug here. Something's not working. I'm falling into an error. I need to fix myself."00:56
Who should watch
You are building agents that may encounter requests or failures your original tool set does not cover, and you want to see a runtime code-generation pattern.
You are evaluating multi-agent systems and need concrete checks for goal completion, tool selection, parameters, call order, and inter-agent messages.
You are considering self-modifying agents and need a practical starting point for sandboxing, permission limits, and execution visibility.