We Let Claude Code and Codex Race Human Researchers

Elie Bakouch, Prime Intellect19:39 · Sept 2026 · 14K views
Thumbnail for We Let Claude Code and Codex Race Human Researchers Watch on YouTube
TL;DR
  1. 1

    The Optimizer Speedrun gives AI agents a measurable task: improve a GPT-2-level training result by changing optimizer-related parameters.

  2. 2

    Claude Code and Codex both beat the human record, but Claude often stopped while Codex kept working and used more scratchpad space, sub-agents, and tokens.

  3. 3

    The agents combined existing research ideas rather than inventing a new optimizer, so Prime Intellect is building a broader discovery loop with generation, evaluation, judging, scaling, and human guidance.

Summary

Elie Bakouch describes an open test of automated AI research using the Optimizer Speedrun, where agents try to train a GPT-2-level model in fewer steps by changing optimizer-related parameters. Claude Code and Codex both beat the human record, though they worked differently. Claude repeatedly stopped and declared the task too hard, while Codex continued with little idle time and used more scratchpad space, sub-agents, and tokens. A longer comparison with Claude, Codex, Kimi, and GLM showed that token efficiency changes the ranking. The agents also used research papers in different ways, and a paper found only by Claude led to the best record. Bakouch expected genuinely new optimizer ideas, but the models only combined existing techniques. Prime Intellect is now working toward a discovery loop inspired by AlphaEvolve, with agents proposing ideas, speedruns providing rewards, judges giving feedback, larger-scale tests, and humans steering the process. He argues that this work should be done in the open.

Key ideas
00:34

Recursive self-improvement needs an independent test

Bakouch frames the project around claims that recursive self-improvement is coming soon. He defines it as models training models without human intervention. There is no benchmark that measures whether this is happening, and no third-party benchmark run outside the large labs. He also wants to understand how models conduct research because future scientific work may depend on AI tools. The project therefore uses a public, repeatable research environment rather than relying on claims from the companies building the models.

01:52

The Optimizer Speedrun turns model research into a constrained race

Andrej Karpathy's GPT-2 speedrun showed that a GPT-2-level validation loss could be reached far faster than the original training process. The community improved the result through the modded-nanogpt project, eventually reaching less than two minutes. The Optimizer Speedrun narrows the allowed changes: researchers can change optimizer-related parameters, such as replacing Adam with another optimizer, while keeping the task focused on finding a better method. Bakouch sees this as more research-oriented than simply rewriting the whole training program for speed.

04:18

A speedrun supplies rules, rewards, and a quick verification loop

Bakouch says speedruns work as evaluation environments because they have a clear target and verifiable results. A successful improvement gets a positive reward, while failure gets zero or a negative reward. Runs take about 15 to 20 minutes, and records must pass a statistical threshold so that a result is not just random noise or an optimization accident. The same setting could also train models because agents receive feedback from a concrete outcome. Its short cycle lets an agent propose an idea, submit a job, inspect the logs, and decide what to try next.

05:32

Claude Code and Codex showed different working habits

Prime Intellect gave Claude Code and Codex access to a goal file and a Slurm cluster with preemptible jobs. The agents could submit experiments, read training logs, and decide whether a result was a record. Claude Code stopped every nine or ten hours and said it could not improve the record, leaving it idle for about a third of the run. Codex kept working, rarely asked questions, and almost never sat idle. Codex also wrote much more into its scratchpad, spawned more sub-agents, and used far more tokens. The difference was behavioral, not simply a matter of one agent running longer.

10:21

Both agents beat the human record

The human record shown in the experiment was about 2,990 steps. Claude Code beat it by roughly 50 or 60 steps, while Codex finished about 20 steps above the human result. The agents could fetch the latest human record during the run, so restarting an agent gave it access to new community progress. Bakouch presents the result as evidence that agents can improve on a strong public baseline in this setting, while also noting that the initial experiment lacked the structure needed for a proper benchmark.

11:36

A serious benchmark needs controlled access tracks

Prime Intellect's planned benchmark uses multiple seeds and puts models under the same conditions. It has three research-access tracks: one where the model relies only on knowledge in its weights, one with access to arXiv papers, and one with full access to current information and the latest human record. The plan covers both the original nanoGPT speedrun and the Optimizer Speedrun, where the allowed changes are constrained. This design separates what models already know from what they can discover through literature search or current external information.

12:45

Token accounting changes which agent looks most efficient

In a longer run lasting almost six days, Claude Code and Codex were effective, and Kimi produced a breakthrough around day four that beat Codex with a new record. Claude improved more progressively, while Kimi's progress looked like a step function. The ranking changed when Bakouch plotted results against output tokens rather than elapsed time. One model consumed far more tokens, while Kimi used its tokens efficiently. The comparison shows why elapsed time alone does not describe the cost or efficiency of automated research.

14:10

The agents combined known methods instead of inventing optimizers

Bakouch expected the agents to produce an optimizer idea that human researchers had not found. They did not. The agents combined ideas from multiple papers and made small improvements on top of existing methods, but none produced a novel optimizer or mechanism. He sees this as a meaningful limitation because the task is difficult while still accessible to human researchers. Claude did find a paper that no other model found, and that paper led to the best record, so literature search helped even without producing a new fundamental method.

15:40

Prime Intellect is building a discovery loop around speedruns

The proposed system is inspired by AlphaEvolve. Multiple generators, including open-source models, suggest ideas. The speedrun evaluates them and produces a reward. A judge gives quality feedback and develops a preference for promising methods. Strong candidates can then be tested with more parameters and tokens because some methods that work in the speedrun may fail at larger scale. Bakouch also gives humans a role in judging ideas and steering agents. Changing the objective and constraints can create many speedruns that push discovery in different directions.

"Codex totally the opposite just worked for all the time and almost never idle never asked for question."08:16
Who should watch
  • You are building agent evaluations and need a research task with a measurable result, short experiment cycles, and explicit access conditions.
  • You want to compare AI research agents on more than their final score, including idle time, token use, scratchpad behavior, sub-agent use, and literature search.
  • You are interested in automated discovery and want to see where current agents stop short of inventing new methods.