Pierce Freeman and Richard Diehl Martinez discuss Andrej Karpathy's autoresearch project, which lets an external coding agent edit a small training program, run experiments and retain changes that improve a validation metric. They connect it with Karpathy's compact educational projects and the value of code that researchers can inspect and modify.
The hosts describe an overnight run that changes training settings, handles failures and uses version control to recover from unsuccessful experiments. They emphasize that a prepared training setup and a capable external agent are doing different jobs: the small model being trained is not independently rewriting its own intelligence.
They compare this feedback loop with coding agents that learn from test results, then ask whether improved small-model efficiency demonstrates frontier research ability. Historical neural architecture search provides a useful comparison, although autoresearch can edit architecture as well as training settings within its allowed file.
The episode ends by distinguishing optimization, novel ideas and the ability to connect discoveries across disciplines. One host revises an initially narrow view of creativity to include effective search. Their disagreement centers on research taste and meaningful novelty, not on whether automated experimentation can already produce useful incremental improvements.
Watch on YouTube




