# How Autoresearch Is Changing ML Research

Zhengyao Jiang, Weco AI | AI Engineer World's Fair 2026 | 16:16

Source: https://www.youtube.com/watch?v=iCj_ATyThvc
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/how-autoresearch-is-changing-ml-research
Published: 2026-07-16
Tags: agents, evals, multi-agent

## TL;DR
- Aiden produced seven leaderboard records in OpenAI's Parameter Golf challenge, while the best human contributor produced three.
- Aiden's strongest results came from finding, combining, and implementing ideas developed by people in papers and community posts.
- Autoresearch shifts human work toward designing evaluations and code abstractions that shape what an agent can discover.

## Summary
Zhengyao Jiang describes Aiden, Weco AI's autonomous research agent, and its 22-day run in OpenAI's Parameter Golf competition. Aiden ran about 1,300 experiments on one H100 node, produced seven leaderboard records, and had a computed H-index of 10 based on community citations. Jiang traces most of its successful ideas to human research papers and suggestions from other competitors. Aiden's contribution was to implement those ideas, find combinations that worked under the 16MB constraint, and execute many experiments quickly. Jiang argues that this is a model for human-AI collaboration: people provide ideas and design the challenge, while agents perform much of the search and execution. He compares autoresearch to training a model. The evaluation is like the loss function, while the code abstraction is like the architecture. Poor abstractions can permit data leakage, while strict interfaces can produce better solutions. Human creativity, judgment, and system design therefore become more valuable as execution is automated.

## Key ideas
### Aiden was built to produce work that a research community could reuse
[01:14](https://www.youtube.com/watch?v=iCj_ATyThvc&t=74s)
Jiang distinguishes Aiden from an agent that only climbs a local benchmark. Weco AI built it to publish its work so other engineers could merge, fork, and build on it. Aiden is a multi-agent, self-improving system that reads public information such as research papers and pull requests, runs experiments, and submits a pull request after its findings pass a quality gate. In the Parameter Golf competition, it ran for about 22 days and produced seven new leaderboard records. The best human contributor produced three. Jiang uses community adoption as a second measure of quality beyond host review and leaderboard scores.

### Aiden had high experimental throughput without using most of the competition's compute
[04:01](https://www.youtube.com/watch?v=iCj_ATyThvc&t=241s)
Over 22 days, Aiden ran about 1,300 experiments on a single H100 node. Jiang says throughput alone does not explain its performance, because the system also kept a high rate of useful submissions. It used at most 4% of the competition's total compute and produced about 15% of its records. About 28% of Aiden's submissions reached the leaderboard, roughly six times the community average. Jiang says this improved the signal-to-noise ratio of the competition's public pull requests. Aiden did not win by massively parallelizing the search, even though autonomous research could support much more parallel work.

### Aiden's strongest results came from executing and combining human ideas
[05:25](https://www.youtube.com/watch?v=iCj_ATyThvc&t=325s)
Jiang traces almost all of Aiden's record pull requests to human research papers, other Parameter Golf participants, or related communities such as nanoGPT. Some ideas were never submitted because people considered implementation difficult. Aiden could find those ideas and carry them through. Jiang gives an example involving gated attention from a Qwen paper. The change added parameters and broke the 16MB file-size limit, so Aiden used quantization to reduce the size. A later contributor posted a tokenizer improvement. Aiden combined it with the architectural changes, and the three ideas produced a large performance jump.

### Execution is often the bottleneck after a useful research idea exists
[07:41](https://www.youtube.com/watch?v=iCj_ATyThvc&t=461s)
Jiang describes autoresearch as strong at finding and implementing ideas, extracting promising ingredients from a noisy public channel, and testing straightforward fixes such as quantization after a parameter increase. Its advantage also comes from searching combinations across a large space quickly. He says these abilities may sound like ordinary execution, but execution is usually the bottleneck. Progress often comes from building on an existing idea and carrying out many good implementations. In his account, the agent does not need to originate most of the concepts to move the competition forward.

### The people who design the competition still shape what the agents can achieve
[09:06](https://www.youtube.com/watch?v=iCj_ATyThvc&t=546s)
Jiang warns against judging the competition only by the engineers doing hill climbing. The design of the challenge strongly affects the value of the community's work. A bad design can make the effort ineffective, while a good design can give autonomous systems a productive direction. He uses an analogy from Andrej Karpathy's earlier comment about gradient descent writing code better than people. As models took over more coding execution, software engineering did not disappear. Higher-level skills became more valuable. Jiang expects autoresearch to create a similar change in ML research and engineering.

### Autoresearch depends on evaluations and code abstractions in the same way models depend on loss and architecture
[11:19](https://www.youtube.com/watch?v=iCj_ATyThvc&t=679s)
Jiang compares running autoresearch with training a model. The evaluation is the signal that trains the code, similar to data and a loss function or an environment in reinforcement learning. It defines what the agent optimizes. The code abstraction provides the framework in which the agent searches, much like a neural network architecture biases which functions are easier to learn. A proprietary evaluation dataset or a strong understanding of what matters in a field can create an advantage. As agents improve, a well-designed evaluation can have a larger effect on their results.

### A strict data-processing API prevented leakage that a loose API allowed
[13:36](https://www.youtube.com/watch?v=iCj_ATyThvc&t=816s)
Jiang describes an autoresearch experiment on a fraud-detection pipeline. The goal was to optimize data preprocessing. With a loose API, the same function processed both training and test data. The resulting score looked good, but test-set information leaked into training, contaminating the solution. Weco then tightened the abstraction so test data could not reach the training process. The data-leakage rate dropped to zero. Jiang uses this example to show that an agent's search space is partly defined by the interfaces engineers give it. A reward can still be hacked if the system permits it, so the abstraction needs careful design.

### Human work moves up the stack toward designing the agent's search space
[14:38](https://www.youtube.com/watch?v=iCj_ATyThvc&t=878s)
Jiang calls autoresearch a new craft: designing a hill for an agent to climb. He expects creativity and judgment in the design of evaluations and abstractions to become more valuable as automated search improves. Driving these systems is itself a new skill that barely existed one or two years earlier. His conclusion is that humans move up the stack rather than leave it. The search becomes automated, while people decide what should be measured, which constraints should apply, and how the agent should explore.

## Notable quotes
- "Can the auto research agent produce work that a human community actually recognize beyond a good score?" (01:14)
- "Aiden was 10 and the next human was seven." (03:41)
- "What moves the frontier is usually exactly some belief on existing ideas and tons of good executions." (08:53)
- "Your eval is the loss function and the data." (11:19)
- "The human would just move up the stack not out of it." (15:10)

## Tools & references mentioned
- Parameter Golf
- Aiden
- Weco AI
- OpenAI
- MLE-bench
- Qwen
- nanoGPT
- Andrej Karpathy

## Who should watch
- You are building an autonomous ML or coding agent and need to decide what it should evaluate and how it should search.
- You want evidence about how an agent can contribute to a public research community instead of only optimizing a private benchmark.
- You design data pipelines, APIs, or competitions where a loose interface could let an automated system exploit the setup.

## Related talks

- [Autoresearch Made Our Models 3x Faster](https://aietalks.com/talks/autoresearch-made-our-models-3x-faster) (Tejas Bhakta, Morph LLM, 07:30)
- [Morgan Stanley's ALPHALAB: Multi-Agent Research Across Optimization Domains](https://aietalks.com/talks/morgan-stanleys-alphalab-multi-agent-research-across-optimization-domains) (Brendan Rappazzo, Morgan Stanley, 20:07)
- [First Steps Toward Automated AI Research](https://aietalks.com/talks/first-steps-toward-automated-ai-research) (Richard Socher, Recursive AI, 20:24)
- [How Anthropic Builds: Lessons from Labs](https://aietalks.com/talks/how-anthropic-builds-lessons-from-labs) (Mike Krieger, Anthropic, 26:11)
- [We Let Claude Code and Codex Race Human Researchers](https://aietalks.com/talks/we-let-claude-code-and-codex-race-human-researchers) (Elie Bakouch, Prime Intellect, 19:39)
