Aman Gupta and Shreya Rajpal present a support-agent evaluation workflow built around simulated customer conversations. They argue that multi-turn agent traces with tool state are slow to author manually, while relying only on production traces limits how quickly teams can test changes without involving live customers.
The simulation method creates customer personas, consistent synthetic account details and multi-turn conversations against a real agent with mocked tools. The resulting traces can be scored with evaluation metrics and used to compare prompts, tools, model choices and agent harness changes before a production experiment.
Aman Gupta reports that Nubank uses this process in an improvement loop: observe agent behavior, generate evaluation data, optimize, catch regressions, and advance the changes that pass. He reports faster iteration and improved customer-service outcomes, including human review indicating many simulated cases were usable; these company results have not been independently verified here.
Shreya Rajpal stresses that simulated data needs a measured link to real outcomes. The closing guidance calls for human review and offline-to-online comparisons to test whether the simulations reflect production behavior, alongside metrics that match the service qualities a team actually wants to improve.
Watch on YouTube




