Can Your Agent Hear You Now? Building Live Voice Agents with Gemini - Thor Schaeff

AI Engineer20m 14s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Thor Schaeff starts with expressive text-to-speech and an illustrated story demo. Gemini Live calls an image-generation tool, while the Interactions API carries forward the previous interaction to maintain story context and visual continuity. The example separates a conversational agent's orchestration from the media model that produces its output.

    Thor Schaeff demonstrates live translation by letting audience members listen to the same talk in their preferred language. A radio-DJ demo then uses a voice conversation to request a song and invoke Lyria 3. These are illustrations of tool-using conversational components rather than claims that every demo directly fits a business need.

    Thor Schaeff explains the Live API as a stateful WebSocket session accepting text, audio and video frames and streaming back audio with a transcript. The native audio model differs from a cascaded speech-recognition, text-model and speech-synthesis pipeline. Interruptions, multilingual interaction and function calls allow a listener to redirect the conversation or connect it to other systems.

    Thor Schaeff finishes with voice and vision in Google AI Studio and describes search grounding as a way to access current information. The important building blocks are session state, native speech understanding, multimodal inputs and explicit tools; conference invitations, subscription requests and unrelated promotion are omitted.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Portraits of Thor Schaeff against black with the headline LIVE VOICE AGENTS in attention blue and white. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 11 October 2026 and duration 20m 14s.

    Thor Schaeff demonstrates how Gemini's native audio, stateful sessions and tool calls support multilingual agents that combine voice, vision and generated media.