Agentic Inference for AI Agents - Byung-Gon Chun

AI Engineer14m 59s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Byung-Gon (Gon) Chun frames an agent's task, rather than an individual model request, as the useful performance unit. Planning, tool calls and observations form a loopAn agent loop repeatedly interprets the current state, chooses an action, observes the result and decides what to do next., sometimes with parallel subagentsA multi-agent system contains multiple AI agents that interact, coordinate, divide work, or influence one another while pursuing tasks., so users experience the time until the whole task finishes.

    Long-running agent steps often reuse a growing context. Prefix cachingPrompt caching lets an AI service reuse previously processed prompt content so repeated context can cost less and run faster. can avoid reprocessing the shared portion, while key-value cache management uses GPU memory, host memory and disk to keep more context available across steps and replicas.

    Cache-aware routing can send a request toward a reusable prefix while balancing load. Agent-aware scheduling can use knowledge of the larger workflow for prefill, preemption and cache eviction decisions, with the aim of reducing end-to-end task latency.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Byung-Gon Chun in a blue sweater gestures beside the blue and white headline FASTER AI AGENTS on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 19 September 2026 and duration 14m 59s.

    Byung-Gon Chun explains how shared-context caching, memory management, routing and scheduling can reduce the time an AI agent takes to finish a task.