Elizabeth Fuentes Leone argues that larger context windowsA context window is the maximum amount of tokenized information an AI model can consider during one processing session. alone do not solve agent reliability. Using an application-log retrieval example, she distinguishes externalizing large data, selecting relevant informationContext engineering designs the information, instructions, memory, and tool state an AI receives so it can perform a task reliably., compressing conversation historyAI context compaction reduces accumulated prompt history while preserving information needed for the agent to continue safely. and isolating context between agents.
The talk walks through Strands Agents conversation-management strategies, then distinguishes short-term conversation history, long-term vector memoryAI agent memory is stored information that an agent can retrieve and use across steps, sessions, or changing contexts. and graph-based relational memory. Memory pointers keep bulky tool results in storage while allowing an agent to retrieve the underlying data when it is actually needed.
For multi-agent workflows, she recommends sharing references rather than flooding every agent with the same context. She also demonstrates explicit tool-call limitsAn agent iteration limit is the maximum number of work-and-revision cycles an AI agent may use for a task. and asynchronous job-status tools to prevent unbounded loops or long external requests from blocking an agent.
Her closing cautions cover context stuffing, lossy summarization and unnecessary context propagation. The practical emphasis is on supplying the information a task requires while retaining clear tool responses and controlled execution.
Watch on YouTube




