Lakshya A Agrawal argues that expensive agent runs contain more useful information than a final reward score alone. GEPA uses tool responses, compiler errors and other textual feedback to propose changes in prompts. He presents task-specific comparisons with reinforcement learningReinforcement learning trains an AI system to choose actions using feedback about the results of its earlier choices. to illustrate possible sample-efficiency gainsSample efficiency measures how much useful learning an AI system achieves from a given number of training examples or experiences., not a claim that prompt optimizationPrompt optimization searches for instructions or examples that improve an AI system's measured performance on a defined task without necessarily changing its model weights. replaces every form of model training.
Lakshya A Agrawal describes a Pareto-style poolA Pareto front is a set of choices where improving one objective requires worsening another, so no included choice is better than another on every measured objective. that retains candidates which succeed on different examples. This diversity helps search escape local optima. The optimize_anything interface extends the approach to scored text artifacts, including code, agent architectures, scheduling policies and numerical parameters serialized as text.
Lakshya A Agrawal presents applications to coding-agent skills, reasoning tasks and subjective evaluation. Human annotations can help optimize an LLM judgeA large language model as a judge is an evaluation method in which a language model scores, compares or critiques another system's output using stated criteria., whose feedback then guides the agent. He closes by discussing work that combines fast prompt adaptation with slower weight updates. The reported improvements depend on the task, evaluator and experimental setting.
Watch on YouTube




