From 36% to 100%: How Self-Improving Agents Write Their Own Skills

Rafal Wilinski, Runlayer18:31 · Oct 2026 · 8,698 views
Thumbnail for From 36% to 100%: How Self-Improving Agents Write Their Own Skills Watch on YouTube
TL;DR
  1. 1

    Skills give agents reusable playbooks that guide long tasks away from costly dead ends.

  2. 2

    MCP can distribute governed skills from one central server to agents and clients across a company.

  3. 3

    Distilling skills from successful and failed runs raised one task's success rate from 36% to 100%.

Summary

Rafal Wilinski argues that agents need reusable procedural knowledge as their tasks become longer. A skill is a playbook that an agent can discover through a short description and load when it becomes relevant. Without one, an agent may spend hours and many tool calls pursuing a bad approach. Wilinski describes three obstacles: skills are usually built for developers, people rarely document hard-won procedures, and many clients do not support them. Runlayer uses MCP as a distribution layer, with a remote server that stores, governs, and serves skills to different agents and clients. The company also distills skills from agent traces. Successful runs provide reusable procedures, while failures expose missing guardrails and edge cases. In one example, a distilled skill improved a Chromium-on-AWS-Lambda task from 36% to 100% success. Wilinski says this creates an organizational flywheel that preserves internal knowledge as models change or become less capable.

Key ideas
00:32

Hard problems should become reusable company playbooks

Wilinski imagines an engineer, salesperson, or lawyer writing down exactly how they solved each hard problem and handing that playbook to the team. In practice, the debugging, reasoning, and procedural knowledge often disappear when a task or session ends. He argues that agents can automate this process, so a solution to one difficult problem becomes available whenever a similar problem appears. Examples include the exact scripts and feature flags for a rollback, or the wording needed when dealing with an enterprise customer who may churn.

02:12

Skills guide agents before long tasks become expensive

A skill is a playbook an agent can read when its knowledge appears relevant. Through progressive disclosure, the client first sees a short description, then loads the full skill into the agent's context when needed. Wilinski connects this to longer task horizons. An agent that can reason for hours may also spend hundreds of tool calls defending a wrong thesis. Skills influence the initial trajectory, helping the agent avoid wasted time, tokens, and money.

05:21

Skills have problems around authorship, safety, and client support

Wilinski identifies three barriers. Existing skills are often markdown files, JSON files, CLI commands, or Git repositories, which makes them accessible to developers but harder to roll out to nontechnical teams. People also rarely have the time and energy to turn a difficult experience into a useful playbook. Finally, support varies across clients, including tools used by marketing, HR, legal, and finance teams. He also warns that installing random skills can introduce prompt injections, root permissions, and exposed API keys.

07:55

MCP can distribute skills from a governed source

Wilinski separates knowledge from capability. Skills describe how to do something, while MCP tools provide capabilities, but MCP can also deliver the skills themselves. A remote MCP server can act as the source of truth for a company. Clients search it for relevant skills, while the server can enforce policy and select the right version based on the query and the user's role. He says Runlayer has used this approach internally and points to the Skills Over MCP working group.

10:10

Voyager turns exploration into reusable procedural knowledge

Wilinski uses the Voyager Minecraft paper as the model for self-improving agents. Voyager did not discard its traces after exploring Minecraft. When it discovered how to accomplish something, it saved the procedure as code containing the tool calls that led to the result. Later, the agent could retrieve that recipe instead of rediscovering how to make a crafting table or diamond pickaxe. Wilinski carries the same idea into company workflows: exploration should produce reusable procedures.

11:34

Failed runs expose what a skill is missing

Runlayer lets agents do real work and asks a Frontier model to distill an interesting successful trace into a skill. Wilinski says failures provide the richest signal because they expose missing guardrails, missing libraries, and edge cases that were not anticipated. Successful runs make a skill more trustworthy, while failed runs show what must be added. The distillation and updates happen asynchronously without a human in the loop.

12:54

A distilled skill raised one task from 36% to 100% success

Runlayer tested an agent on running a Chromium fork inside AWS Lambda. The task initially succeeded 36% of the time. After one agent invocation solved the problem, the skill distiller created a reusable skill and included it in future runs. The success ratio reached 100%. Wilinski says the benefit is also a smaller search space, with fewer branches, tool calls, tokens, and costs.

14:04

A company-wide skill library can preserve knowledge as models change

Wilinski says frontier intelligence is rented because models can be deprecated or become less capable. Companies should use access to strong models to discover and preserve useful procedures. Runlayer combines skills, MCP, and a central knowledge base that receives new data. Similar runs are grouped, distilled skills pass checks for prompt injection and PII, and approved skills are served through the MCP gateway to every client and agent. This creates a company-wide record of hard-won lessons.

15:23

The organizational flywheel can become internal knowledge competitors cannot copy

When one agent solves a hard problem, the solution can become a governed skill available to other agents and people facing the same situation. Wilinski calls this a self-improving organizational flywheel. He presents the accumulated procedures as a possible source of differentiation when models and software become widely available. The flywheel raises the floor for everyone in the company, preserves internal practices, and can lower costs without changing the model.

"A skill to me is a playbook that essentially agent can read when it thinks that the knowledge is relevant."02:23
Who should watch
  • You are building agents that run for a long time and need them to avoid repeating expensive mistakes.
  • Your company has useful operational knowledge spread across people, sessions, and departments, but no reliable way to turn it into reusable agent instructions.
  • You are deciding whether skills should be installed locally or delivered through a governed MCP service.