Thor Schaeff starts with expressive text-to-speech and an illustrated story demo. Gemini Live calls an image-generation tool, while the Interactions API carries forward the previous interaction to maintain story context and visual continuity. The example separates a conversational agent's orchestration from the media model that produces its output.
Thor Schaeff demonstrates live translation by letting audience members listen to the same talk in their preferred language. A radio-DJ demo then uses a voice conversation to request a song and invoke Lyria 3. These are illustrations of tool-using conversational components rather than claims that every demo directly fits a business need.
Thor Schaeff explains the Live API as a stateful WebSocket session accepting text, audio and video frames and streaming back audio with a transcript. The native audio model differs from a cascaded speech-recognition, text-model and speech-synthesis pipeline. Interruptions, multilingual interaction and function calls allow a listener to redirect the conversation or connect it to other systems.
Thor Schaeff finishes with voice and vision in Google AI Studio and describes search grounding as a way to access current information. The important building blocks are session state, native speech understanding, multimodal inputs and explicit tools; conference invitations, subscription requests and unrelated promotion are omitted.
Watch on YouTube




