Doug Guthrie builds an observability workflow around a support agent using the OpenAI Agents SDK and Braintrust. He starts with tracing inputs, outputs and tool calls, then distinguishes deterministic checks, LLM judges and human calibration as sources of quality signals.
Doug Guthrie configures online scoring and topic analysis to find patterns beyond known test cases. The workshop covers custom facets, preprocessing, sampling, conversation grouping and querying trace metadata, using imported support traces to investigate record-lookup failures and other recurring issues.
Doug Guthrie demonstrates bringing selected production examples back into evaluation datasets and using a coding agent with the Braintrust CLI to investigate failures, propose changes and compare evaluation results. He emphasizes a reviewable improvement cycle rather than treating collected traces as useful on their own.
The questions explore deployment models, access restrictions, scoring cost, changing test sets and remote evaluations. Doug Guthrie shows how parameterized evaluations can expose an agent to a playground while keeping code execution on a separate server.
Watch on YouTube




