Google AI Edge engineer Cormac Brick contrasts cloud inference with local intelligence that can work offline, preserve data on the device and provide predictable latency. Memory cost and the wider range of target hardware constrain deployment: fitting model weights is only part of the budget once the runtime, cache and operating system are included.
Brick compares general small models with much smaller specialists for speech, vision, embeddings and function calling. He reports hardware-dependent performance and demonstrates why merely running a model is not enough if an interaction feels too slow. The transcript's contradictory tiny-model parameter phrase is not adopted as a technical specification.
Task-specific fine-tuning with synthetic data is presented as a route to compact, responsive features. Examples include voice commands mapped to a fixed set of functions and offline dictation using separate speech-recognition and text-polishing models. Brick treats easier customization and faster visual perception as future work; reported reliability remains specific to the demonstrated tasks.
Watch on YouTube




