Rustem Feyzkhanov argues that each organization needs benchmarks reflecting its own tools, policies and workflows. He contrasts production traces with repeatable offline experiments that compare success, cost, latency and retries while holding the environment and evaluators steady.
Rustem Feyzkhanov describes executable tasks with realistic environments, oracle solutions and verifiers that examine final state, traces and artifacts. He covers benchmark CI, reward hacking, changing one component at a time, expanding tests from production failures and keeping an unseen holdout. Audience questions address common and edge-case coverage, an illustrative 80/20 split and expert review where evaluators disagree. Scale and performance claims are attributed to the speaker.
Watch on YouTube




