Tim Scarfe interviews Alexander Mattick about what different learning formulations actually optimize. They compare probabilistic and energy-based models with diffusion and flow-based approaches, focusing on sampling costs, normalization and how training can amortize expensive computation. Alexander Mattick distinguishes theoretical expressiveness from practical efficiency rather than declaring a universal winner.
They discuss mathematical accounts of deep learning and the need for theories that make testable predictions. The conversation then challenges a purely reward-driven view of learning: acquiring information through exploration can be expensive, and known structure should not always be relearned from scratch.
Alexander Mattick explains why explicit constraints can make reinforcement-learning objectives easier to specify than repeatedly tuning a single weighted reward. Expected-cost limits, tail-risk objectives and action restrictions address different problems; none automatically guarantees worst-case safety in every environment.
The final discussion examines the overloaded term world model. Predicting a system is not the same as planning reliably within it, and simulated rollouts do not remove difficulties with physical contact, data coverage and deployment latency. The speakers distinguish polished demonstrations from dependable autonomy. The sponsor segment is omitted.
Watch on YouTube




