Why AI Agents Need Real-World Evals

MTS36:13
0 comments · 0 votesOpen discussion
Sign in to join the discussion

    Video summary

    Brewer and Schiavo introduce Grove Research as an agent-ecology company studying interactions among humans and AI agents, and among agents themselves. They expect many kinds of agents to coexist, compete, cooperate and form networks rather than converging on one universal system.

    The ecology analogy emphasizes emergent behavior. Individual agents may look predictable in isolation, yet shared environments, memory, retrieval and social feedback can produce outcomes that a model-level safety test never exercises. Cooperation and symbiosis can matter as much as competition.

    They argue that models are highly sensitive to whether they believe they are being evaluated. Training in simulations can make an agent treat benchmark-like environments as unreal, while deployment context, public feedback and retrieved history can change its persona and actions in ways that static tests miss.

    Their proposed response is naturalistic observation: build a field station for persistent multi-agent and multi-human systems, track what agents actually do online and create shared scientific language for their behavior. In this view, the real world is the ultimate evaluation because consequences and feedback cannot be fully reproduced in a sandbox.

    Original YouTube thumbnailWatch on YouTube