Building Closed-Loop Evals for a Multimodal Agent at Scale: Soumya Gupta and Jai Chopra, Uber

AI Engineer21m 39s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Soumya Gupta and Jai Chopra describe a multimodal pipeline that routes merchant photographs for enhancement or leaves them untouched. Its goals include better quality while preserving ingredients, portions and each merchant's identity, rather than making every photograph look the same.

    Human-labelled datasets spanning dishes, regions and image quality calibrate routing precision and recall. Production samples reveal drift, and diagnostic, reflection and synthesis agents propose configuration changes that must pass existing benchmarks before promotion.

    Image-specific editing prompts receive limited quality-feedback retries. Pairwise checks assess faithfulness, completeness and realism: invented shrimp, missing sauce and superficial overcorrection remain failures. A separate publication check can catch errors missed upstream.

    The presenters combine human reviews, merchant feedback and production outcomes to target weak components. Their approach prioritises observability and bounded, benchmarked updates rather than an unconstrained self-improving loop.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Soumya Gupta and Jai Chopra beside the blue-and-white headline “CLOSED-LOOP AI EVALS” on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 24 July 2026 and duration 21m 39s.

    Uber's food-photo agent uses human-calibrated routing, bounded editing and layered quality checks to improve images without changing the dish.