GEPA uses full rollout traces and textual reflection to improve prompts with far fewer examples than reinforcement learning.
2
A Pareto pool helps GEPA avoid the local optima that trap a simple iterative prompt-improvement loop.
3
Optimize Anything applies the same method to prompts, code, agent harnesses, skills, scheduling policies, and other text-defined systems that can be scored.
Summary
Lakshya A. Agrawal presents GEPA as a sample-efficient alternative to reinforcement learning for improving AI systems. Reinforcement learning reduces a rollout to a final score, while GEPA gives a language model the full trace, including reasoning, tool calls, responses, and error messages. The model uses that information to write a better prompt, while a Pareto pool preserves candidates that perform well on different examples. Agrawal reports large gains from very small amounts of reflection, including improvements to GPT-4.1 mini, an AMD NPU coding agent, and several benchmarks. Optimize Anything extends the method beyond prompts to Python programs, agent harnesses, coding skills, kernels, and scheduling policies. Examples include improving Gemini Flash on ARC-AGI from 32.5% to 89.5% and a GPT-5 mini coding agent from 24% to 93%. The talk also covers production use, learned evaluators, and joint optimization of prompts and model weights.
Reinforcement learning discards most of the information in an agent rollout
Agrawal says current reinforcement learning often runs many rollouts, waits for a final reward, and turns that reward into gradients. A rollout also contains chains of thought, tool calls, environment responses, and error messages. Those details can explain why the system failed, but a method such as GRPO reduces them to an O(1) score before learning from them. This creates a sample-efficiency problem when agents run for hours, use long-running tools, or rely on expensive task metrics.
GEPA reflects on full traces and changes prompts in natural language
GEPA uses a language model or agent to read the full trace of a rollout and describe what worked and what failed. The reflection can use intermediate outputs, retrieval from a company knowledge base, documentation, or other tools. GEPA then updates a prompt rather than making only small weight changes. Agrawal compares this with changing a summarization instruction from 'generate a one-line summary' to 'generate a 10-line summary'. A small textual change can produce a large behavior change.
Three examples can outperform thousands of GRPO rollouts in the reported comparison
Agrawal compares GEPA with GRPO on a training domain. After one reflection round using three data points, GEPA reaches twice the performance gains that GRPO reaches after 25,000 rollouts. More GEPA steps widen the gap again. The example uses Qwen3 8B optimizing itself, without an external expert teacher. GEPA produces detailed task specifications, including how to interpret inputs, the purpose of a pipeline stage, and lessons extracted from the data.
GEPA can discover missing domain instructions in a prompt
For AMD's XDNA2 NPU, the programming API had very little information available online and GPT-4 was failing on the task. An existing agent scored 4.25%, and GEPA raised it to 30.52% without changing the agent apart from its prompt. One discovered instruction was to avoid including ADF.h. AMD ships that library for NPU programming, but it did not work with the generation of hardware in the example. GEPA found this in one step.
The Pareto pool prevents prompt search from settling on one local optimum
GEPA runs the system on examples, collects domain-specific feedback, asks a model to propose a better prompt, and keeps candidates in a Pareto pool. The pool retains every candidate that wins on at least one training example instead of keeping only the highest overall scorer. Agrawal's comparison shows a simple loop repeatedly refining one candidate until it gets stuck. The Pareto search keeps a more balanced set of alternatives and reaches a higher score. Across four benchmarks, this accounted for more than half of GEPA's gains.
Optimize Anything applies reflective search to any text-defined, scored artifact
Agrawal extends GEPA beyond prompts because an agent harness can be written as Python or JavaScript, and other systems can also be serialized as text. Optimize Anything takes problems, an evaluator or fitness function, and any available side information. That information can include compiler errors, profiler output, tool-call failures, documentation, job traces, or SLA violations. The evaluator scores each candidate, while an LLM proposes improved versions and the Pareto pool preserves useful alternatives.
Reflective search can discover complete agent architectures
Starting from a four-line Python program that asked a model to solve ARC-AGI through chain of thought, Optimize Anything found a six-step agent in 16 reflection rounds. Gemini Flash's ARC-AGI accuracy rose from 32.5% to 89.5%. The discovered agent performs rule and hypothesis induction, synthesizes code, executes and traces that code, debugs it, proposes new versions, and runs the final code on test inputs. On MATH-500, a two-step agent improved GPT-4.1 nano's accuracy by 20%.
Cheap optimization of skills can transfer to a stronger coding model
Agrawal describes optimizing repository-specific skills for a coding agent. Using a GPT-5 mini agent under a tight budget, the method raised Go issue-resolution performance from 24% to 93%. The learned skill contained information about repository structure, test commands, feature locations, and the build system. Applying the skill to Claude Sonnet 4.5 reached 100% issue resolution in the reported example and cut resolution time roughly in half.
Production workflows can learn evaluators and jointly optimize prompts and weights
For subjective tasks, teams can collect production traces and have a human annotate about 50 trajectories with detailed judgments. GEPA can optimize an LLM-as-a-judge prompt from those annotations, then use that evaluator to improve the agent. Agrawal describes this as a continuing data flywheel. He also mentions fast-slow learning, which co-optimizes model weights with prompt and agent harness changes. Other examples include lower OCR error rates and a Databricks deployment that tuned GPT-OSS 120B to outperform Claude Opus at 90 times lower cost.
"Our idea is to perform reflective optimization in text space where instead of only using the zero or one reward signal, we can have a language model or an agent look at the trace of the entire rollout and reflect on what worked in them, what did not work in them."03:02
Who should watch
You are improving an agent with expensive rollouts and have useful traces, tool errors, or domain documentation that a final reward currently ignores.
Your team spends time hand-tuning prompts, tool descriptions, control flow, or repository-specific coding skills and wants to search over those artifacts automatically.
You need to optimize code, an agent harness, a policy, or another text-defined system with an evaluator but have too little data for conventional fine-tuning.