Why 80% Reliability Isn't Good Enough - Felipe Blanes, Amazon AGI Lab

AI Engineer17m 44s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Felipe Blanes describes the gap between strong benchmark results and the unexpected tasks customers attempt in production. He proposes an evaluation flywheel: define success from the customer's perspective, capture signals, diagnose gaps and feed prioritized findings into product, engineering and research decisions.

    The loop combines instrumentation with direct customer conversations because metrics alone can misrepresent the problem. Early adopters reveal possible uses, broader deployment exposes common workflows, and later feedback identifies product gaps. Evaluations must change as those needs become more complex.

    Felipe Blanes argues that unreliable agents can add supervision and recovery work rather than save time. He recommends being explicit about what works, what is improving and what remains out of scope. Examples include caching successful trajectories with model fallback and accommodating different levels of technical skill, reinforcing a practical focus on production-derived scenarios rather than imagined success.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Felipe Blanes in blue rests his hand at his chin beside the blue and white "EARN TRUST" headline. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 9 October 2026 and duration 17m 44s.

    Agent evaluations should evolve from real customer needs and production failures, rather than treating a static benchmark score as proof of trustworthiness.