Scaling to Long Horizons - Ross Taylor and Chengxi Taylor

AI Engineer18m 7s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Ross Taylor contrasts Galactica's base-model demo with ChatGPT's reinforcement learning from human feedback, arguing that a strong base model alone did not make a useful product. He recounts work on curated data, repeated training, thinking tokens and a Llama 2 mathematics-and-reinforcement-learning recipe, then attributes later reflective reasoning to stronger base models, larger context windows and more reinforcement learning compute.

    Chengxi Taylor frames long-horizon tasks as a problem of limited context, sparse rewards, credit assignment and variable-length trajectories. She describes context compaction, value models, file-system scratchpads and search over prior work as ways to carry useful information across a long task, while noting the extra complexity and possible bias that value models introduce.

    For evaluation, Chengxi Taylor cites a year-long football prediction benchmark in which, she reports, the tested frontier models lost their starting budgets. She argues that common tests give too little attention to open-ended, interactive tasks. She then explains how pipelined reinforcement learning can improve GPU use while increasing off-policy drift, and how bootstrapping before a task ends exchanges idle time for value-model bias.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Ross Taylor and Chengxi Taylor in blue tops beside the blue and white headline Long-Horizon Agents on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 31 July 2026 and duration 18m 7s.

    Ross Taylor and Chengxi Taylor argue that long-horizon agents need stronger base models, durable memory, better credit assignment, realistic evaluations and careful compute trade-offs.