Aparna Dhinakaran explains why a fixed grading rubric can miss failures in a multi-step agent workflow. Tool use, loops, changing interfaces and intermediate decisions all affect the outcome, so evaluation needs to examine the trajectory and its context rather than score only a final answer.
She proposes agent-based judges as a complement to deterministic checks and LLM-based scoring, not a replacement for either. Production traces can reveal recurring tool inefficiencies and other behavioral patterns, giving engineers evidence for proposed fixes. The talk presents a development direction rather than claiming that adaptive judges eliminate evaluation risk.
Watch on YouTube




