From Agent Traces to Agent Simulations

AI Engineer20m 24s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Rustem Feyzkhanov argues that each organization needs benchmarks reflecting its own tools, policies and workflows. He contrasts production traces with repeatable offline experiments that compare success, cost, latency and retries while holding the environment and evaluators steady.

    Rustem Feyzkhanov describes executable tasks with realistic environments, oracle solutions and verifiers that examine final state, traces and artifacts. He covers benchmark CI, reward hacking, changing one component at a time, expanding tests from production failures and keeping an unseen holdout. Audience questions address common and edge-case coverage, an illustrative 80/20 split and expert review where evaluators disagree. Scale and performance claims are attributed to the speaker.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Rustem Feyzkhanov with one hand resting at his chin beside the blue and white headline TEST THE WHOLE AGENT on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 25 July 2026 and duration 20m 24s.

    Rustem Feyzkhanov explains how to turn production traces into repeatable simulations that test an AI agent's entire stack, not just its model.