Hossein Niazmandi distinguishes evaluations before production from observability after deployment. He explains why nondeterministic model behavior makes both necessary, then decomposes a basic evaluation into test inputs, execution outputs and scores. His example illustrates how curated tests can help compare versions without proving an agent will perform reliably on real traffic.
Hossein Niazmandi traces a progression from spreadsheets through custom interfaces and prompt experiments to production traces. The useful feedback loop turns observed failures into reproducible test cases, measures proposed changes and watches deployed behavior for new regressions. Engineers, product managers and domain specialists need different interfaces to participate in that process.
Hossein Niazmandi describes the engineering burden behind the full system: large nested traces, full-text search, aggregate analytics, asynchronous scoring, annotations and access controls. He considers agents as additional users of evaluation tools while retaining human judgment over improvements. Braintrust supplies his examples, but the talk's central systems challenges apply beyond any single vendor.
Watch on YouTube




