Laurie Voss explains why nondeterministic agents require more than ordinary code inspection. Nested traces expose prompts, tool calls, outputs, latency and cost; evaluations add scores and explanations so teams can find recurring failure patterns without manually reading every interaction.
Laurie Voss works through a toy-store shopping agent with observability instrumentation and trace-analysis skills. A broken price filter produces empty results, and the workshop demonstrates diagnosing and checking a correction. The demo also encounters setup problems and uses an existing solution, so it is an illustration of the workflow rather than proof of autonomous debugging success.
Laurie Voss describes periodically examining production traces to propose issues, evaluations and code changes. Humans remain responsible for accepting or rejecting those suggestions, while regression evaluations protect previously working behavior and guardrails test domain boundaries. He also discusses deployment and compliance limitations of the demonstrated tooling.
Watch on YouTube




