Tejas Kumar distinguishes evaluation before release from protections applied by an agent harness at runtime. He introduces test cases, expected behavior, judges and aggregation, then discusses position, verbosity, self-preference and sycophancy as reasons an automated judge can give misleading results.
The live coding example starts with a refund-policy scenario, moves beyond a brittle text check and compares judge decisions with a labeled dataset. Kumar adds policy context, changes model when necessary and builds a CI agreement gate. He closes with retrieval-backed policy and fresh production examples; his numerical thresholds are choices for the demonstration rather than universal reliability guarantees.
Watch on YouTube




