Aiden produced seven leaderboard records in OpenAI's Parameter Golf challenge, while the best human contributor produced three.
2
Aiden's strongest results came from finding, combining, and implementing ideas developed by people in papers and community posts.
3
Autoresearch shifts human work toward designing evaluations and code abstractions that shape what an agent can discover.
Summary
Zhengyao Jiang describes Aiden, Weco AI's autonomous research agent, and its 22-day run in OpenAI's Parameter Golf competition. Aiden ran about 1,300 experiments on one H100 node, produced seven leaderboard records, and had a computed H-index of 10 based on community citations. Jiang traces most of its successful ideas to human research papers and suggestions from other competitors. Aiden's contribution was to implement those ideas, find combinations that worked under the 16MB constraint, and execute many experiments quickly. Jiang argues that this is a model for human-AI collaboration: people provide ideas and design the challenge, while agents perform much of the search and execution. He compares autoresearch to training a model. The evaluation is like the loss function, while the code abstraction is like the architecture. Poor abstractions can permit data leakage, while strict interfaces can produce better solutions. Human creativity, judgment, and system design therefore become more valuable as execution is automated.
Aiden was built to produce work that a research community could reuse
Jiang distinguishes Aiden from an agent that only climbs a local benchmark. Weco AI built it to publish its work so other engineers could merge, fork, and build on it. Aiden is a multi-agent, self-improving system that reads public information such as research papers and pull requests, runs experiments, and submits a pull request after its findings pass a quality gate. In the Parameter Golf competition, it ran for about 22 days and produced seven new leaderboard records. The best human contributor produced three. Jiang uses community adoption as a second measure of quality beyond host review and leaderboard scores.
Aiden had high experimental throughput without using most of the competition's compute
Over 22 days, Aiden ran about 1,300 experiments on a single H100 node. Jiang says throughput alone does not explain its performance, because the system also kept a high rate of useful submissions. It used at most 4% of the competition's total compute and produced about 15% of its records. About 28% of Aiden's submissions reached the leaderboard, roughly six times the community average. Jiang says this improved the signal-to-noise ratio of the competition's public pull requests. Aiden did not win by massively parallelizing the search, even though autonomous research could support much more parallel work.
Aiden's strongest results came from executing and combining human ideas
Jiang traces almost all of Aiden's record pull requests to human research papers, other Parameter Golf participants, or related communities such as nanoGPT. Some ideas were never submitted because people considered implementation difficult. Aiden could find those ideas and carry them through. Jiang gives an example involving gated attention from a Qwen paper. The change added parameters and broke the 16MB file-size limit, so Aiden used quantization to reduce the size. A later contributor posted a tokenizer improvement. Aiden combined it with the architectural changes, and the three ideas produced a large performance jump.
Execution is often the bottleneck after a useful research idea exists
Jiang describes autoresearch as strong at finding and implementing ideas, extracting promising ingredients from a noisy public channel, and testing straightforward fixes such as quantization after a parameter increase. Its advantage also comes from searching combinations across a large space quickly. He says these abilities may sound like ordinary execution, but execution is usually the bottleneck. Progress often comes from building on an existing idea and carrying out many good implementations. In his account, the agent does not need to originate most of the concepts to move the competition forward.
The people who design the competition still shape what the agents can achieve
Jiang warns against judging the competition only by the engineers doing hill climbing. The design of the challenge strongly affects the value of the community's work. A bad design can make the effort ineffective, while a good design can give autonomous systems a productive direction. He uses an analogy from Andrej Karpathy's earlier comment about gradient descent writing code better than people. As models took over more coding execution, software engineering did not disappear. Higher-level skills became more valuable. Jiang expects autoresearch to create a similar change in ML research and engineering.
Autoresearch depends on evaluations and code abstractions in the same way models depend on loss and architecture
Jiang compares running autoresearch with training a model. The evaluation is the signal that trains the code, similar to data and a loss function or an environment in reinforcement learning. It defines what the agent optimizes. The code abstraction provides the framework in which the agent searches, much like a neural network architecture biases which functions are easier to learn. A proprietary evaluation dataset or a strong understanding of what matters in a field can create an advantage. As agents improve, a well-designed evaluation can have a larger effect on their results.
A strict data-processing API prevented leakage that a loose API allowed
Jiang describes an autoresearch experiment on a fraud-detection pipeline. The goal was to optimize data preprocessing. With a loose API, the same function processed both training and test data. The resulting score looked good, but test-set information leaked into training, contaminating the solution. Weco then tightened the abstraction so test data could not reach the training process. The data-leakage rate dropped to zero. Jiang uses this example to show that an agent's search space is partly defined by the interfaces engineers give it. A reward can still be hacked if the system permits it, so the abstraction needs careful design.
Human work moves up the stack toward designing the agent's search space
Jiang calls autoresearch a new craft: designing a hill for an agent to climb. He expects creativity and judgment in the design of evaluations and abstractions to become more valuable as automated search improves. Driving these systems is itself a new skill that barely existed one or two years earlier. His conclusion is that humans move up the stack rather than leave it. The search becomes automated, while people decide what should be measured, which constraints should apply, and how the agent should explore.