Felipe Blanes describes the gap between strong benchmark resultsA benchmark is a standardized task or collection of tests used to compare AI systems under defined conditions. and the unexpected tasks customers attempt in production. He proposes an evaluation flywheelAgent evaluation tests whether an AI agent completes tasks correctly, consistently, and within its required boundaries.: define success from the customer's perspective, capture signals, diagnose gaps and feed prioritized findings into product, engineering and research decisions.
The loop combines instrumentation with direct customer conversations because metrics alone can misrepresent the problem. Early adopters reveal possible uses, broader deployment exposes common workflows, and later feedback identifies product gaps. Evaluations must change as those needs become more complex.
Felipe Blanes argues that unreliable agentsAI agent reliability is the degree to which an agent consistently completes intended work correctly and safely under expected conditions. can add supervision and recovery work rather than save time. He recommends being explicit about what works, what is improving and what remains out of scope. Examples include caching successful trajectories with model fallback and accommodating different levels of technical skill, reinforcing a practical focus on production-derived scenarios rather than imagined success.
Watch on YouTube




