Tim Scarfe interviews Shawn Wen about the time-sensitive mechanics of voice conversations and the constraints enterprise customers place on assistants. Shawn Wen describes an audio-native model that predicts turn-taking signals before producing text responses, citations and a later transcript for auditing.
Tim Scarfe and Shawn Wen discuss noisy call data, privacy handling, synthetic augmentation and latency-aware reasoning. Shawn Wen explains why over-cleaned audio can hurt the model and why caller expectations, regional voices and backend delays complicate otherwise impressive demonstrations.
Tim Scarfe and Shawn Wen examine evaluation beyond isolated speech-recognition scores, including response time, tool use and whether customers actually finish their tasks. The later discussion covers enterprise harness ownership, cognitive debt and the shift from generating work to reviewing it. The episode was produced in partnership with PolyAI; product-performance claims remain attributed to the speaker.
Watch on YouTube




