Fuad Ali uses a shopping assistant that returns products above a customer's budget to explain why launch-time evaluations miss some production failures. Instrumentation records what the agent actually did; pairing those traces with repository context lets an investigator connect the wrong output to a faulty price-filter code path.
Fuad Ali describes two routes for that investigation. Arize Signal continuously groups recurring failures, collects evidence and proposes tickets or pull requests. Arize skills let a developer's coding agent query traces, curate datasets and run evaluations on demand. Both approaches use operational evidence alongside code rather than relying on an isolated dashboard or an agent's confidence.
Fuad Ali distinguishes a fixed LLM judge rubric from a tool-using harness that can inspect the trajectory and external context. Access checks, valid tool arguments and a user's requested price range can require information beyond the input and final response. The budget failure becomes an explicit evaluation of whether every returned product satisfies the constraint.
Fuad Ali closes the loop with experiments that replay actual failure cases against a development endpoint. Comparing the proposed fix with production behavior can reveal changes in correctness, latency and cost before release. The workshop retains human code review and presents verification on real failures as a necessary step, not proof that autonomous repair is always safe.
Watch on YouTube




