From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 - Laurie Voss, Arize AI

AI Engineer42m 17s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Laurie Voss explains why nondeterministic agents require more than ordinary code inspection. Nested traces expose prompts, tool calls, outputs, latency and cost; evaluations add scores and explanations so teams can find recurring failure patterns without manually reading every interaction.

    Laurie Voss works through a toy-store shopping agent with observability instrumentation and trace-analysis skills. A broken price filter produces empty results, and the workshop demonstrates diagnosing and checking a correction. The demo also encounters setup problems and uses an existing solution, so it is an illustration of the workflow rather than proof of autonomous debugging success.

    Laurie Voss describes periodically examining production traces to propose issues, evaluations and code changes. Humans remain responsible for accepting or rejecting those suggestions, while regression evaluations protect previously working behavior and guardrails test domain boundaries. He also discusses deployment and compliance limitations of the demonstrated tooling.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Laurie Voss against a black background beside the blue-and-white headline “FIX AGENTS WITH FEEDBACK”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 5 October 2026 and duration 42m 17s.

    Laurie Voss shows how traces, evaluations and human-approved changes can turn real agent failures into a repeatable improvement and regression-testing loop.