Harbor Rollouts: Agent Evals, Production and Training

AI Engineer21m 11s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Alex Shaw treats agent development as empirical machine learning rather than predictable conventional software. Model choice, instructions, tools and environments interact probabilistically, making repeated trials and relevant evaluations essential. The session, credited by the publisher to Shaw and Ryan Marten, uses Harbor to give those trials a shared execution format.

    Harbor combines instructions, a sandboxed environment and a verifier, preserving the agent trajectory and the resulting reward. Shaw recommends building evaluations around real work before choosing a model, then reusing the same machinery for parallel production tasks. A map-and-reduce example extracts recurring human corrections from coding sessions to inform better tests.

    The resulting trajectories can also support supervised training, reinforcement learning and iterative skill improvement. These applications share infrastructure, but their reliability still depends on task design, useful verification and attention to overfitting or reward hacking. Hiring invitations and benchmark-tour promotion are omitted from this account.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Alex Shaw and Ryan Marten against a black background beside the blue and white headline “AGENT ROLLOUTS EVAL TO TRAINING”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 24 July 2026 and duration 21m 11s.

    In a session credited to Alex Shaw and Ryan Marten, Shaw explains how Harbor connects agent evaluation, production work and training through sandboxed tasks, preserved trajectories and verifiable outcomes.