How DoorDash Makes AI Evals a Team Sport

AI Engineer16:11
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Paranjape and Chitlur Haridas describe DoorDash's GenAI Platform team as a horizontal group that helps product teams balance accuracy, latency and cost. Its shared primitives include LLM and agent gateways, open-weight model hosting and a common evaluation platform.

    Their central lesson is that AI quality depends on domain knowledge held across strategy and operations, product, labeling partners and engineering. The platform therefore began with approachable user interfaces, expanded into reusable APIs and now supports workflow-first access so each group can contribute without waiting on the central team.

    The continuous loop begins with production traces and sessions, samples them into manageable review sets, adds domain-specific annotations and turns the reviewed examples into golden datasets. Teams then calibrate LLM-as-judge prompts against those datasets, inspect the changes and monitor the resulting scores over time.

    A telemetry surface exposes traces, scores and observations through APIs, SDKs and MCP, while a workflow surface supports annotation, dataset review and judge calibration. Self-serve annotation tools and transparent prompt comparisons reduced annotation costs, shortened feedback cycles and let different teams choose who owns each quality decision.

    Original YouTube thumbnailWatch on YouTube