Building an AI Observability and Evaluation Workflow

AI Engineer1h 51m
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Doug Guthrie builds an observability workflow around a support agent using the OpenAI Agents SDK and Braintrust. He starts with tracing inputs, outputs and tool calls, then distinguishes deterministic checks, LLM judges and human calibration as sources of quality signals.

    Doug Guthrie configures online scoring and topic analysis to find patterns beyond known test cases. The workshop covers custom facets, preprocessing, sampling, conversation grouping and querying trace metadata, using imported support traces to investigate record-lookup failures and other recurring issues.

    Doug Guthrie demonstrates bringing selected production examples back into evaluation datasets and using a coding agent with the Braintrust CLI to investigate failures, propose changes and compare evaluation results. He emphasizes a reviewable improvement cycle rather than treating collected traces as useful on their own.

    The questions explore deployment models, access restrictions, scoring cost, changing test sets and remote evaluations. Doug Guthrie shows how parameterized evaluations can expose an agent to a playground while keeping code execution on a separate server.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Doug Guthrie beside the headline Observe Your AI on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 6 October 2026 and duration 1h 51m.

    Doug Guthrie demonstrates how production traces, evaluators and topic clustering can feed a tested improvement cycle for AI agents.