Alex Shaw treats agent development as empirical machine learning rather than predictable conventional software. Model choice, instructions, tools and environments interact probabilistically, making repeated trials and relevant evaluations essential. The session, credited by the publisher to Shaw and Ryan Marten, uses Harbor to give those trials a shared execution format.
Harbor combines instructions, a sandboxed environment and a verifier, preserving the agent trajectory and the resulting reward. Shaw recommends building evaluations around real work before choosing a model, then reusing the same machinery for parallel production tasks. A map-and-reduce example extracts recurring human corrections from coding sessions to inform better tests.
The resulting trajectories can also support supervised training, reinforcement learning and iterative skill improvement. These applications share infrastructure, but their reliability still depends on task design, useful verification and attention to overfitting or reward hacking. Hiring invitations and benchmark-tour promotion are omitted from this account.
Watch on YouTube




