Zubin Aysola explains how the ARIA team at Weights & Biases uses the same agent code in production and simulation. Shared tracingAI agent observability makes an agent's state, actions, tool use, failures, resource use, and outcomes visible enough to understand and operate it. makes it possible to reproduce useful or problematic interactions as offline tasks, while regular synchronization reduces drift between the environments.
Zubin Aysola describes configurable experiments with environment setup, agent execution, scoring and teardown. Evaluations include task completionAgent evaluation tests whether an AI agent completes tasks correctly, consistently, and within its required boundaries. and relative behavioral preferences, using both simple instructions and simulated multi-turn users. The team examines trajectoriesAI trajectory evaluation assesses the sequence of reasoning-relevant states, tool calls, decisions, and side effects that led to an agent's final result. rather than relying on a single aggregate score.
Zubin Aysola's live walkthrough asks ARIA to convert a production trace into a regression taskAI regression testing reruns preserved cases to detect whether a model or agent update has broken behavior that previously worked. and compare a candidate with the production agentPrompt optimization searches for instructions or examples that improve an AI system's measured performance on a defined task without necessarily changing its model weights.. The reported change addresses an SDK logging error through a small prompt or skill adjustment. He stresses that automating this work does not remove the need for human judgment about evaluation quality, guardrails and which improvements matter.
Watch on YouTube




