Operating Distributed Inference Systems at Scale: Nishant Gupta and Naman Ahuja

AI Engineer19m 51s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Nishant Gupta explains why large AI inference workloadsAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action. differ from ordinary web services. Agent-based applications can make many model calls per user task, request sizes vary sharply, and GPU capacity, batching and KV cache statePrompt caching lets an AI service reuse previously processed prompt content so repeated context can cost less and run faster. make each placement decision consequential. He traces a request through gateways, routing, caching, scheduling and model serving, arguing that failures and latency can arise anywhere along that path.

    Gupta proposes scheduling that accounts for GPU type, memory headroom, loaded model weights, cache state, tenant priority, workflow progress and latency budget. His optimization framework asks whether work can be avoided, shared, moved or delayed; he recommends evaluating cost per successful taskCost per completed task measures the total AI, tool, infrastructure, retry, and repair expense for each verified useful outcome. rather than optimizing token throughputAI inference throughput measures how much model-serving work a system completes per unit of time. alone. He also explains how retries and cold GPU pools can amplify a local degradation into a wider failure, motivating circuit breakers, admission control, load shedding and retry budgets.

    Naman Ahuja closes by treating inference as a distributed systems problem with control loopsAn AI feedback loop occurs when AI outputs or their consequences become inputs that influence later model or agent behavior. fed by first-token time, utilization, success per dollar and request-path latency. He describes the tradeoff among latency, throughput and cost, then argues that routing, batching, caching, scheduling and reliability are converging into an inference control plane.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Nishant Gupta in a cap and sunglasses and Naman Ahuja in a knit beanie flank the blue-and-white INFERENCE AT SCALE headline and a GPU routing diagram on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 19 September 2026 and duration 19m 51s.

    Nishant Gupta and Naman Ahuja argue that scaling AI inference depends on workload-aware scheduling, reliable control loops and measuring cost per successful task.