How We Built an Agent That Improves Itself

AI Engineer17m 6s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Zubin Aysola explains how the ARIA team at Weights & Biases uses the same agent code in production and simulation. Shared tracingAI agent observability makes an agent's state, actions, tool use, failures, resource use, and outcomes visible enough to understand and operate it. makes it possible to reproduce useful or problematic interactions as offline tasks, while regular synchronization reduces drift between the environments.

    Zubin Aysola describes configurable experiments with environment setup, agent execution, scoring and teardown. Evaluations include task completionAgent evaluation tests whether an AI agent completes tasks correctly, consistently, and within its required boundaries. and relative behavioral preferences, using both simple instructions and simulated multi-turn users. The team examines trajectoriesAI trajectory evaluation assesses the sequence of reasoning-relevant states, tool calls, decisions, and side effects that led to an agent's final result. rather than relying on a single aggregate score.

    Zubin Aysola's live walkthrough asks ARIA to convert a production trace into a regression taskAI regression testing reruns preserved cases to detect whether a model or agent update has broken behavior that previously worked. and compare a candidate with the production agentPrompt optimization searches for instructions or examples that improve an AI system's measured performance on a defined task without necessarily changing its model weights.. The reported change addresses an SDK logging error through a small prompt or skill adjustment. He stresses that automating this work does not remove the need for human judgment about evaluation quality, guardrails and which improvements matter.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Zubin Aysola in blue and glasses, gesturing beside “ARIA LEARNS FROM EVALS” in blue and white on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 26 September 2026 and duration 17m 6s.

    Zubin Aysola demonstrates a feedback loop that turns real agent interactions into offline evaluation tasks. ARIA can propose and test changes to its own prompts or skills, while people remain responsible for choosing useful improvements.