Pim de Witte explains General Intuition's approach to learning visual policies from gameplay and other action-labeled video. The interview connects game-controller inputs with interfaces robots already support, aiming to transfer learned behavior without replacing existing locomotion and balancing systems.
Pim de Witte distinguishes policy learning from predicting pixels alone: visible outcomes may omit the hidden actions that caused them. Human action trajectories therefore remain important, even when bot-generated data can teach some environment dynamics. Multi-view world-model training is discussed as a way to encourage shared spatial structure.
Pim de Witte says transfer becomes harder as embodiment and action spaces become more complex, including humanoid hands not represented in current training data. Forecasts about simulation-driven growth and convergence between code and pixel-based intelligence are presented as his expectations, not established results.
Watch on YouTube




