Adam Lucek on Building Agent Evals You Can Hill Climb

Adam Lucek22m 53s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Adam Lucek distinguishes regression checks from benchmarks designed to measure and improve a specific agent capability. He breaks a task into harness, state, verifier and configuration, then walks through a legal-document benchmark to show how instructions, workspace tools, source documents and explicit pass/fail criteria fit together. The example illustrates evaluation design rather than offering legal advice.

    Adam Lucek emphasizes the difficult work of defining a capability, curating representative inputs and labeling desired outcomes with domain experts. He recommends environments that reproduce production conditions and warns that agents noticing incomplete mocks can undermine an experiment. Once those foundations are reliable, teams can compare changes to models, prompts and harnesses or use verifier scores as reinforcement-learning rewards. Promotional appeals are omitted, and industry examples are presented as his explanation rather than independently audited performance findings.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Adam Łucek beside the blue and white headline Evals You Can Climb on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 4 October 2026 and duration 22m 53s.

    Adam Lucek explains how curated datasets, realistic environments and trustworthy verifiers turn agent evaluations into useful optimization targets.