Can Rewriting an AI Agent Bend the Intelligence Curve?

Machine Learning Street Talk43m 43s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Zhengyao Jiang joins Tim Scarfe to discuss Weco's experiment in letting one research agent modify the harnessAn AI agent harness is the software framework that packages a model with tools, instructions, context management, execution controls, and user interaction. around another. Zhengyao Jiang reports that eight days of automated experimentation exceeded results from two years of manual engineering on held-out tasksAn evaluation set is a collection of examples kept for measuring an AI system rather than training it.. The model weights stayed fixed: the changes involved code, prompts, tools and the way experiments were organized.

    Zhengyao Jiang describes evaluation splits, new benchmark families and safeguards against reward hackingReward hacking happens when an AI system exploits a scoring rule or proxy to earn a high reward without achieving the intended outcome.. One statistical detector itself stopped working during optimization, illustrating why a higher score is not enough to establish a reliable advance. He distinguishes repairing a broken evaluator from exploiting it and argues that tests must evolve alongside agent capabilities.

    Zhengyao Jiang separates delegation and useful research output from a stronger claim: an optimizer becoming better at improving itselfRecursive self-improvement is the proposed process in which an AI system helps improve its own capabilities, then uses those improvements to support further advances.. He says this experiment did not establish that stronger level. Tim Scarfe and Zhengyao Jiang also discuss memoryAI agent memory is stored information that an agent can retrieve and use across steps, sessions, or changing contexts., search budgets, open-ended exploration and the continued importance of human-created abstractions and research directions.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Tim Scarfe in blue on the left and Zhengyao Jiang in white on the right, framing “AGENTS REWRITING AGENTS” in blue and white on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 26 September 2026 and duration 43m 43s.

    Zhengyao Jiang explains how optimizing an agent's code, prompts and tools can improve research performance without changing its underlying model, while distinguishing task gains from a system that improves its own ability to improve.