Why Building an Eval Platform Is Harder Than It Looks - Braintrust

AI Engineer17m 47s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Hossein Niazmandi distinguishes evaluations before production from observability after deployment. He explains why nondeterministic model behavior makes both necessary, then decomposes a basic evaluation into test inputs, execution outputs and scores. His example illustrates how curated tests can help compare versions without proving an agent will perform reliably on real traffic.

    Hossein Niazmandi traces a progression from spreadsheets through custom interfaces and prompt experiments to production traces. The useful feedback loop turns observed failures into reproducible test cases, measures proposed changes and watches deployed behavior for new regressions. Engineers, product managers and domain specialists need different interfaces to participate in that process.

    Hossein Niazmandi describes the engineering burden behind the full system: large nested traces, full-text search, aggregate analytics, asynchronous scoring, annotations and access controls. He considers agents as additional users of evaluation tools while retaining human judgment over improvements. Braintrust supplies his examples, but the talk's central systems challenges apply beyond any single vendor.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Hossein Niazmandi beside the blue and white headline “AI EVALUATIONS - BEYOND SPREADSHEETS” on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 6 October 2026 and duration 17m 47s.

    Hossein Niazmandi argues that reliable agent evaluation requires a continuous production-feedback loop and substantial data infrastructure, not merely a spreadsheet interface.