The speakers describe failures of LLM judges in open-ended browser tasks, where changing websites and several valid completion paths make deterministic checks difficult. Their verifier builds a task-specific rubric, retrieves relevant screenshots for each criterion and compares observed outcomes with the agent's claims.
The design separates partial process credit from final task success, avoids adding requirements the user never requested and distinguishes controllable mistakes from environmental obstacles. Examples show why hallucinated numbers, cascading penalties and out-of-stock products need different treatment.
They report human-label agreement and training-data filtering experiments, then compare manual verifier development with automated research loops. The strongest reported approach combines human findings with automated experimentation. A closing discussion describes held-out annotations and possible extensions to desktop workflows; these results are presented as the speakers' research findings, not independent replication.
Watch on YouTube




