AI Evals: Judge Calibration, Policy and CI Gates

AI Engineer59m 49s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Tejas Kumar distinguishes evaluation before release from protections applied by an agent harness at runtime. He introduces test cases, expected behavior, judges and aggregation, then discusses position, verbosity, self-preference and sycophancy as reasons an automated judge can give misleading results.

    The live coding example starts with a refund-policy scenario, moves beyond a brittle text check and compares judge decisions with a labeled dataset. Kumar adds policy context, changes model when necessary and builds a CI agreement gate. He closes with retrieval-backed policy and fresh production examples; his numerical thresholds are choices for the demonstration rather than universal reliability guarantees.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Tejas Kumar against a black background beside the blue and white headline “CALIBRATE AI JUDGES”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 5 October 2026 and duration 59m 49s.

    Tejas Kumar builds a policy-grounded LLM judge and an agreement-based CI gate, showing why biased judges, missing context and stale data can undermine apparently passing evaluations.