Routing LLM Inference in Production: Qianru Lao and Lu Zhang

AI Engineer18m 12s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Lu Zhang describes an inference load balancerLoad balancing distributes requests across available workers or servers so an AI service can use capacity and manage waiting times. that selects a GPU-backed engine for each request while accounting for engine capability, response latency, reliability, geographic distance and cached conversation contextPrompt caching lets an AI service reuse previously processed prompt content so repeated context can cost less and run faster.. An earlier weighted hashing scheme adjusted engine weights from live performance signals, but the feedback loopAn AI feedback loop occurs when AI outputs or their consequences become inputs that influence later model or agent behavior. could be hard to explain and could oscillate traffic between engines, weakening cache reuse.

    Qianru Lao outlines a newer split between a control plane with a global view of traffic and engine capacity and a data plane that makes each request decision from a local routing snapshot. The control plane combines live signals with capacity and latency profiles to publish routing weights asynchronously. A regional example shows why sending a request to a farther engine can be faster when a nearby engine is overloaded, because end-to-end latency includes both network time and engine waiting time.

    Lu Zhang closes with safeguards for imperfect production conditions: reducing traffic to degraded engines, limiting retries to avoid retry storms and shedding load when demand exceeds capacity. The talk presents these as practical defenses around a routing policy, rather than a guarantee that every request can be served under failure.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Qianru Lao on the left and Lu Zhang on the right in blue tops beside the white and blue headline ROUTING INFERENCE on a black background. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 19 September 2026 and duration 18m 12s.

    Lu Zhang and Qianru Lao explain how OpenAI moved from reactive engine weights to globally planned inference routing while keeping fast local decisions and overload safeguards.