Paige Bailey introduces the Gemma model family and demonstrates ways to run small models close to the user rather than relying on a hosted service for every interaction. Browser and mobile examples cover text, images, audio and tool use, with deployment choices tied to device memory and available compute.
The session connects these examples to developer tooling, including a browser storybook built with Transformers.js and a mobile model gallery. Bailey describes model sizes, quantization and local execution options, while positioning broader DeepMind scientific work as context rather than claiming that every research capability is demonstrated on a phone.
Watch on YouTube




