Tim Scarfe and Shawn Wen on Voice-Agent Turn Taking

Machine Learning Street Talk1h 10m
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Tim Scarfe interviews Shawn Wen about the time-sensitive mechanics of voice conversations and the constraints enterprise customers place on assistants. Shawn Wen describes an audio-native model that predicts turn-taking signals before producing text responses, citations and a later transcript for auditing.

    Tim Scarfe and Shawn Wen discuss noisy call data, privacy handling, synthetic augmentation and latency-aware reasoning. Shawn Wen explains why over-cleaned audio can hurt the model and why caller expectations, regional voices and backend delays complicate otherwise impressive demonstrations.

    Tim Scarfe and Shawn Wen examine evaluation beyond isolated speech-recognition scores, including response time, tool use and whether customers actually finish their tasks. The later discussion covers enterprise harness ownership, cognitive debt and the shift from generating work to reviewing it. The episode was produced in partnership with PolyAI; product-performance claims remain attributed to the speaker.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Tim Scarfe, Shawn Wen beside the blue and white headline VOICE AGENT TIMING on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 1 October 2026 and duration 1h 10m.

    Tim Scarfe and Shawn Wen explain why reliable voice agents need adaptive turn taking, low latency and auditable outputs rather than a simple speech-to-text pipeline.