Gurvir Singh and Rayan Garg distinguish human time horizons from model measures such as tokens, steps and tool calls. Human completion time depends on expertise and methodology, while model measures vary with the model and harness. They argue that these views should be considered together because neither alone establishes how difficult a task is for an agent.
The talk separates tool coordination, changing environment state and ambiguity as dimensions of task complexity. Simply chaining unrelated tasks makes a trajectory longer without necessarily testing deeper capability. More meaningful tasks require early decisions to influence later work, and leave enough uncertainty for agents to explore rather than follow one prescribed solution.
Gurvir Singh and Rayan Garg describe judges as agents that may need to inspect both the final environment and the path taken to reach it. A deployment-repair example shows why a judge must check logs and outcomes, not just accept tool-call claims. The proposed safeguards include read-only access for judges, scrutiny of reward hacking, and queryable records of long trajectories.
The closing discussion connects rubric design to learnability and evaluation quality. It favors combining deterministic checks with model judges, allowing multiple valid solutions, testing agreement with experts, and examining the breadth of task coverage. The presenters criticize narrow or saturated finance benchmarks and use company data as an illustration; those comparisons are presented as their own assessment rather than independently verified results.
Watch on YouTube




