Beating RL With Reflection: GEPA and Optimize Anything

AI Engineer21m 27s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Lakshya A Agrawal argues that expensive agent runs contain more useful information than a final reward score alone. GEPA uses tool responses, compiler errors and other textual feedback to propose changes in prompts. He presents task-specific comparisons with reinforcement learningReinforcement learning trains an AI system to choose actions using feedback about the results of its earlier choices. to illustrate possible sample-efficiency gainsSample efficiency measures how much useful learning an AI system achieves from a given number of training examples or experiences., not a claim that prompt optimizationPrompt optimization searches for instructions or examples that improve an AI system's measured performance on a defined task without necessarily changing its model weights. replaces every form of model training.

    Lakshya A Agrawal describes a Pareto-style poolA Pareto front is a set of choices where improving one objective requires worsening another, so no included choice is better than another on every measured objective. that retains candidates which succeed on different examples. This diversity helps search escape local optima. The optimize_anything interface extends the approach to scored text artifacts, including code, agent architectures, scheduling policies and numerical parameters serialized as text.

    Lakshya A Agrawal presents applications to coding-agent skills, reasoning tasks and subjective evaluation. Human annotations can help optimize an LLM judgeA large language model as a judge is an evaluation method in which a language model scores, compares or critiques another system's output using stated criteria., whose feedback then guides the agent. He closes by discussing work that combines fast prompt adaptation with slower weight updates. The reported improvements depend on the task, evaluator and experimental setting.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Lakshya A Agrawal in blue and glasses, gesturing beside the blue and white headline “GEPA OPTIMIZE ANYTHING” on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 26 September 2026 and duration 21m 27s.

    Lakshya A Agrawal explains reflective optimization: using execution traces and domain feedback to improve prompts and agent code. GEPA preserves diverse promising candidates instead of repeatedly refining only the highest-scoring one.