Long-Horizon Agents Need Experiments, Not Just Prompts

AI Engineer21m 27s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Erina Karati introduces Project Paradox, a modular framework developed at Supercell's AI Innovation Lab for game agents with individual memoriesAI agent memory is stored information that an agent can retrieve and use across steps, sessions, or changing contexts., emotions, beliefs and plans. Short interactions worked well, but longer runs exposed failuresA long-horizon agent pursues an objective across many actions or extended periods, requiring reliable task state, feedback and stopping conditions.: rumors became facts, sources disappeared and remembered information did not reliably influence actions.

    Erina Karati proposes an external experiment loop that runs controlled scenarios, records traces and changes only a small policy surface. A balanced scorecard measuresAgent evaluation tests whether an AI agent completes tasks correctly, consistently, and within its required boundaries. information spread, source retentionData provenance records where information came from, how it changed, and which people, systems, or processes handled it., uncertainty, action consistency and privacy. Optimizing just one score can reward oversharing or stale memories, so a change is retained only when the relevant guardrails also hold.

    Erina Karati emphasizes freezing the harness, scenarios and metrics while testing memory, retrievalRetrieval is the process of selecting relevant stored information and returning it to an AI system for the current task. and communication policies. A reported improvement in one rumor scenario does not establish general improvement across the system. She presents the approach as an engineering pattern that may also help support, research and workflow agents maintain reliable state over time.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Erina Karati in blue on the right, gesturing beside the blue and white headline “AGENTS NEED EXPERIMENTS” on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 26 September 2026 and duration 21m 27s.

    Erina Karati argues that long-running agents need controlled scenarios and multi-dimensional evaluation, not just memory and better prompts. Project Paradox provides a test environment for studying source retention, uncertainty and planning.