Ross Taylor contrasts Galactica's base-model demo with ChatGPT's reinforcement learning from human feedback, arguing that a strong base model alone did not make a useful product. He recounts work on curated data, repeated training, thinking tokens and a Llama 2 mathematics-and-reinforcement-learning recipe, then attributes later reflective reasoning to stronger base models, larger context windows and more reinforcement learning compute.
Chengxi Taylor frames long-horizon tasks as a problem of limited context, sparse rewards, credit assignment and variable-length trajectories. She describes context compaction, value models, file-system scratchpads and search over prior work as ways to carry useful information across a long task, while noting the extra complexity and possible bias that value models introduce.
For evaluation, Chengxi Taylor cites a year-long football prediction benchmark in which, she reports, the tested frontier models lost their starting budgets. She argues that common tests give too little attention to open-ended, interactive tasks. She then explains how pipelined reinforcement learning can improve GPU use while increasing off-policy drift, and how bootstrapping before a task ends exchanges idle time for value-model bias.
Watch on YouTube




