Felipe Blanes describes the gap between strong benchmark results and the unexpected tasks customers attempt in production. He proposes an evaluation flywheel: define success from the customer's perspective, capture signals, diagnose gaps and feed prioritized findings into product, engineering and research decisions.
The loop combines instrumentation with direct customer conversations because metrics alone can misrepresent the problem. Early adopters reveal possible uses, broader deployment exposes common workflows, and later feedback identifies product gaps. Evaluations must change as those needs become more complex.
Felipe Blanes argues that unreliable agents can add supervision and recovery work rather than save time. He recommends being explicit about what works, what is improving and what remains out of scope. Examples include caching successful trajectories with model fallback and accommodating different levels of technical skill, reinforcing a practical focus on production-derived scenarios rather than imagined success.
Watch on YouTube




