Roland Gavrilescu separates an agent's working loopAn agent loop repeatedly interprets the current state, chooses an action, observes the result and decides what to do next. from a second loop that studies its outputs and improves the system. The strength of the input signal and the verifier determines whether progress is meaningful; recurring failures can become evaluation casesAgent evaluation tests whether an AI agent completes tasks correctly, consistently, and within its required boundaries., repeated actions can become skillsAn AI agent skill is a reusable package of instructions, resources, and tool guidance for performing a bounded kind of work., and user frustration can inform harness changesAn AI agent harness is the software framework that packages a model with tools, instructions, context management, execution controls, and user interaction..
Roland Gavrilescu proposes portable, versioned agent recipes combining prompts, tools, evaluations, model choices and operational context. The point is to preserve how a team arrived at its judgment of good behavior, rather than tying its product to a particular provider or treating the model alone as its advantage.
Roland Gavrilescu uses a recruiting agent to illustrate the process: notice an unwanted pattern, ask a human to calibrate the desired behavior, let agents build tests and candidate changesAI regression testing reruns preserved cases to detect whether a model or agent update has broken behavior that previously worked., then check whether real users prefer the result. The final criterion is valuable work relative to compute cost, after usefulness has been established, not merely an improved offline score.
Watch on YouTube




