From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 - Laurie Voss, Arize AI

AI Engineer1h 51m
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Laurie Voss explains why agent evaluation needs both observability and tests. Traces reveal tool calls, intermediate decisions and failures that final answers can hide, while code-based checks and model judges serve different purposes. Creative but valid solutions should not fail merely because they take an unexpected path.

    Laurie Voss demonstrates a financial-reporting agent and uses its traces to identify problems before writing evaluators. He distinguishes correctness from faithfulness to retrieved context, develops explicit actionability rubrics and calibrates model judges against human-labelled examples. Random demonstration labels and small sample results are not evidence of validated production accuracy.

    Laurie Voss turns observed failures into regression datasets and compares revised prompts through controlled experiments. He recommends improving data and instructions before changing models or hyperparameters, then carrying the same feedback loop into sampled production monitoring. The approach treats evaluation as continuing evidence gathering rather than a one-time guarantee.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Laurie Voss against a black background beside the blue-and-white headline “EVALUATE AGENTS NOT VIBES”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 5 October 2026 and duration 1h 51m.

    Laurie Voss builds a practical agent-evaluation workflow around traces, human-reviewed failure cases, calibrated judges and regression experiments.