# We Let Claude Code and Codex Race Human Researchers

Elie Bakouch, Prime Intellect | AI Engineer | 19:39

Source: https://www.youtube.com/watch?v=oVsEddfhdxc
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/we-let-claude-code-and-codex-race-human-researchers
Published: 2026-09-26
Tags: agents, benchmarks, cost

## TL;DR
- The Optimizer Speedrun gives AI agents a measurable task: improve a GPT-2-level training result by changing optimizer-related parameters.
- Claude Code and Codex both beat the human record, but Claude often stopped while Codex kept working and used more scratchpad space, sub-agents, and tokens.
- The agents combined existing research ideas rather than inventing a new optimizer, so Prime Intellect is building a broader discovery loop with generation, evaluation, judging, scaling, and human guidance.

## Summary
Elie Bakouch describes an open test of automated AI research using the Optimizer Speedrun, where agents try to train a GPT-2-level model in fewer steps by changing optimizer-related parameters. Claude Code and Codex both beat the human record, though they worked differently. Claude repeatedly stopped and declared the task too hard, while Codex continued with little idle time and used more scratchpad space, sub-agents, and tokens. A longer comparison with Claude, Codex, Kimi, and GLM showed that token efficiency changes the ranking. The agents also used research papers in different ways, and a paper found only by Claude led to the best record. Bakouch expected genuinely new optimizer ideas, but the models only combined existing techniques. Prime Intellect is now working toward a discovery loop inspired by AlphaEvolve, with agents proposing ideas, speedruns providing rewards, judges giving feedback, larger-scale tests, and humans steering the process. He argues that this work should be done in the open.

## Key ideas
### Recursive self-improvement needs an independent test
[00:34](https://www.youtube.com/watch?v=oVsEddfhdxc&t=34s)
Bakouch frames the project around claims that recursive self-improvement is coming soon. He defines it as models training models without human intervention. There is no benchmark that measures whether this is happening, and no third-party benchmark run outside the large labs. He also wants to understand how models conduct research because future scientific work may depend on AI tools. The project therefore uses a public, repeatable research environment rather than relying on claims from the companies building the models.

### The Optimizer Speedrun turns model research into a constrained race
[01:52](https://www.youtube.com/watch?v=oVsEddfhdxc&t=112s)
Andrej Karpathy's GPT-2 speedrun showed that a GPT-2-level validation loss could be reached far faster than the original training process. The community improved the result through the modded-nanogpt project, eventually reaching less than two minutes. The Optimizer Speedrun narrows the allowed changes: researchers can change optimizer-related parameters, such as replacing Adam with another optimizer, while keeping the task focused on finding a better method. Bakouch sees this as more research-oriented than simply rewriting the whole training program for speed.

### A speedrun supplies rules, rewards, and a quick verification loop
[04:18](https://www.youtube.com/watch?v=oVsEddfhdxc&t=258s)
Bakouch says speedruns work as evaluation environments because they have a clear target and verifiable results. A successful improvement gets a positive reward, while failure gets zero or a negative reward. Runs take about 15 to 20 minutes, and records must pass a statistical threshold so that a result is not just random noise or an optimization accident. The same setting could also train models because agents receive feedback from a concrete outcome. Its short cycle lets an agent propose an idea, submit a job, inspect the logs, and decide what to try next.

### Claude Code and Codex showed different working habits
[05:32](https://www.youtube.com/watch?v=oVsEddfhdxc&t=332s)
Prime Intellect gave Claude Code and Codex access to a goal file and a Slurm cluster with preemptible jobs. The agents could submit experiments, read training logs, and decide whether a result was a record. Claude Code stopped every nine or ten hours and said it could not improve the record, leaving it idle for about a third of the run. Codex kept working, rarely asked questions, and almost never sat idle. Codex also wrote much more into its scratchpad, spawned more sub-agents, and used far more tokens. The difference was behavioral, not simply a matter of one agent running longer.

### Both agents beat the human record
[10:21](https://www.youtube.com/watch?v=oVsEddfhdxc&t=621s)
The human record shown in the experiment was about 2,990 steps. Claude Code beat it by roughly 50 or 60 steps, while Codex finished about 20 steps above the human result. The agents could fetch the latest human record during the run, so restarting an agent gave it access to new community progress. Bakouch presents the result as evidence that agents can improve on a strong public baseline in this setting, while also noting that the initial experiment lacked the structure needed for a proper benchmark.

### A serious benchmark needs controlled access tracks
[11:36](https://www.youtube.com/watch?v=oVsEddfhdxc&t=696s)
Prime Intellect's planned benchmark uses multiple seeds and puts models under the same conditions. It has three research-access tracks: one where the model relies only on knowledge in its weights, one with access to arXiv papers, and one with full access to current information and the latest human record. The plan covers both the original nanoGPT speedrun and the Optimizer Speedrun, where the allowed changes are constrained. This design separates what models already know from what they can discover through literature search or current external information.

### Token accounting changes which agent looks most efficient
[12:45](https://www.youtube.com/watch?v=oVsEddfhdxc&t=765s)
In a longer run lasting almost six days, Claude Code and Codex were effective, and Kimi produced a breakthrough around day four that beat Codex with a new record. Claude improved more progressively, while Kimi's progress looked like a step function. The ranking changed when Bakouch plotted results against output tokens rather than elapsed time. One model consumed far more tokens, while Kimi used its tokens efficiently. The comparison shows why elapsed time alone does not describe the cost or efficiency of automated research.

### The agents combined known methods instead of inventing optimizers
[14:10](https://www.youtube.com/watch?v=oVsEddfhdxc&t=850s)
Bakouch expected the agents to produce an optimizer idea that human researchers had not found. They did not. The agents combined ideas from multiple papers and made small improvements on top of existing methods, but none produced a novel optimizer or mechanism. He sees this as a meaningful limitation because the task is difficult while still accessible to human researchers. Claude did find a paper that no other model found, and that paper led to the best record, so literature search helped even without producing a new fundamental method.

### Prime Intellect is building a discovery loop around speedruns
[15:40](https://www.youtube.com/watch?v=oVsEddfhdxc&t=940s)
The proposed system is inspired by AlphaEvolve. Multiple generators, including open-source models, suggest ideas. The speedrun evaluates them and produces a reward. A judge gives quality feedback and develops a preference for promising methods. Strong candidates can then be tested with more parameters and tokens because some methods that work in the speedrun may fail at larger scale. Bakouch also gives humans a role in judging ideas and steering agents. Changing the objective and constraints can create many speedruns that push discovery in different directions.

## Notable quotes
- "We don't have any benchmark to basically quantify if this is true or not." (00:34)
- "Onethird of the time the Claude Code agent was idle because I had no way to basically monitor it." (07:52)
- "Codex totally the opposite just worked for all the time and almost never idle never asked for question." (08:16)
- "There was really like no novel optimizer or mechanism that was coming from those model." (14:35)
- "It's super important to have a part of recursive self-improvement to happen in the open." (18:54)

## Tools & references mentioned
- Prime Intellect
- Claude Code
- Codex
- Andrej Karpathy
- GPT-2
- modded-nanogpt
- Keller Jordan
- Slurm
- Adam
- Kimi
- GLM
- AlphaEvolve
- Google
- arXiv

## Who should watch
- You are building agent evaluations and need a research task with a measurable result, short experiment cycles, and explicit access conditions.
- You want to compare AI research agents on more than their final score, including idle time, token use, scratchpad behavior, sub-agent use, and literature search.
- You are interested in automated discovery and want to see where current agents stop short of inventing new methods.

## Related talks

- [Finetuning: 500M AI Agents in Production with 2 Engineers](https://aietalks.com/talks/finetuning-500m-ai-agents-in-production-with-2-engineers) (Mustafa Ali, Method Financial & Kyle Corbitt, OpenPipe, 18:44)
- [How Google DeepMind Runs Agents at Scale](https://aietalks.com/talks/how-google-deepmind-runs-agents-at-scale) (KP Sawhney & Ian Ballantyne, Google DeepMind, 25:13)
- [Why Agent Hype Can Fall Short of Reality](https://aietalks.com/talks/why-agent-hype-can-fall-short-of-reality) (Joel Becker, METR, 21:22)
- [Beating RL With Reflection: GEPA and Optimize Anything](https://aietalks.com/talks/beating-rl-with-reflection-gepa-and-optimize-anything) (Lakshya Agrawal, GEPA, 21:28)
- [The Art & Science of Benchmarking Agents](https://aietalks.com/talks/the-art-science-of-benchmarking-agents) (Vincent Chen, Snorkel AI, 23:25)
