Byung-Gon (Gon) Chun frames an agent's task, rather than an individual model request, as the useful performance unit. Planning, tool calls and observations form a loopAn agent loop repeatedly interprets the current state, chooses an action, observes the result and decides what to do next., sometimes with parallel subagentsA multi-agent system contains multiple AI agents that interact, coordinate, divide work, or influence one another while pursuing tasks., so users experience the time until the whole task finishes.
Long-running agent steps often reuse a growing context. Prefix cachingPrompt caching lets an AI service reuse previously processed prompt content so repeated context can cost less and run faster. can avoid reprocessing the shared portion, while key-value cache management uses GPU memory, host memory and disk to keep more context available across steps and replicas.
Cache-aware routing can send a request toward a reusable prefix while balancing load. Agent-aware scheduling can use knowledge of the larger workflow for prefill, preemption and cache eviction decisions, with the aim of reducing end-to-end task latency.
Watch on YouTube




