Laurie Voss explains why agent evaluation needs both observability and tests. Traces reveal tool calls, intermediate decisions and failures that final answers can hide, while code-based checks and model judges serve different purposes. Creative but valid solutions should not fail merely because they take an unexpected path.
Laurie Voss demonstrates a financial-reporting agent and uses its traces to identify problems before writing evaluators. He distinguishes correctness from faithfulness to retrieved context, develops explicit actionability rubrics and calibrates model judges against human-labelled examples. Random demonstration labels and small sample results are not evidence of validated production accuracy.
Laurie Voss turns observed failures into regression datasets and compares revised prompts through controlled experiments. He recommends improving data and instructions before changing models or hyperparameters, then carrying the same feedback loop into sampled production monitoring. The approach treats evaluation as continuing evidence gathering rather than a one-time guarantee.
Watch on YouTube




