Olivio Sarikas opens Kindle on an Android phone, shares the screen with Gemini and asks for a book suited to a trip in Italy. The assistant recommends The Leopard and points to its cover in the visible collectionVisual grounding connects words in an AI request or answer to the specific objects or regions visible in an image or screen., illustrating how a voice conversation can refer to what is on screenA multimodal model can process or generate more than one kind of data, such as text, images and audio..
Olivio Sarikas then asks about a page introducing formal logic. Gemini explains validity as a relationship between premises and conclusion, and answers his objection that real-world conditions could make the rain example misleading. The exchange shows the value of asking follow-up questions while reading, while keeping the assistant's explanation distinct from a verified assessment of the textbook.
In a second book, Olivio Sarikas asks Gemini to identify Hayao Miyazaki's Starting Point 1979 to 1996, summarize the visible page, read part of it aloudText-to-speech converts written text into spoken audio by predicting pronunciation, timing, prosody and a waveform or intermediate audio representation. and respond to a request to read in Italian. The Italian output is not transcribed, so the transcript supports the request and transition but not an assessment of translation quality.
The final demonstration asks Gemini to create detailed notes in Google KeepApplication integration connects software systems so data and actions can move through one useful workflow. about the page and a separate list of Hayao Miyazaki films. The transcript records the assistant's confirmation and Olivio Sarikas's description of the notes, illustrating a workflow from on-screen reading to saved reference material.
Watch on YouTube




