Varun Singh uses the provocative claim that the base model is dead to question an older training recipe built mainly around broad web text. He argues that reasoning and software-agent tasks give reinforcement learning a larger role, so the base model must provide useful starting skills for that later stage rather than only general knowledge.
Varun Singh contrasts recent data recipes that avoid model-generated text with others that introduce instruction-like and synthetic material during pre-training. He describes rephrasing selected source items into multiple forms as one way to increase data variety and expose the model to the shape of downstream tasks before post-training begins.
Varun Singh also discusses why a sharp change in data distribution between pre-training and post-training can complicate mixture-of-experts routing. He considers mid-training and longer-context agent traces as bridges, then frames supervised learning as a way to teach component skills that reinforcement learning can combine. He presents the balance between these stages as an evolving research question, not a settled formula.
Watch on YouTube




