The demonstration asks an Apptronik Apollo humanoid to pack sports equipment for two children by using calendar context, identifying the right objects and navigating a cluttered shelf. A second challenge asks the robot to find a tall bag on the floor and place it on a table, testing whether it can react to a changed scene.
Jie Tan explains that the embodied reasoning model interprets the world and natural-language instruction, then calls a vision-language-action model to produce motor actions. The robot must coordinate every joint and actuator from feet to fingertips while adjusting its legs in fractions of a second to stay balanced.
The system is designed to recognize failed attempts and try again, which matters for long, physical workflows where one error can invalidate later steps. Google DeepMind presents whole-body control as a necessary foundation for more general robots that can complete useful tasks in real environments. Standard channel calls to action are omitted.
Watch on YouTube


