Lu Zhang describes an inference load balancerLoad balancing distributes requests across available workers or servers so an AI service can use capacity and manage waiting times. that selects a GPU-backed engine for each request while accounting for engine capability, response latency, reliability, geographic distance and cached conversation contextPrompt caching lets an AI service reuse previously processed prompt content so repeated context can cost less and run faster.. An earlier weighted hashing scheme adjusted engine weights from live performance signals, but the feedback loopAn AI feedback loop occurs when AI outputs or their consequences become inputs that influence later model or agent behavior. could be hard to explain and could oscillate traffic between engines, weakening cache reuse.
Qianru Lao outlines a newer split between a control plane with a global view of traffic and engine capacity and a data plane that makes each request decision from a local routing snapshot. The control plane combines live signals with capacity and latency profiles to publish routing weights asynchronously. A regional example shows why sending a request to a farther engine can be faster when a nearby engine is overloaded, because end-to-end latency includes both network time and engine waiting time.
Lu Zhang closes with safeguards for imperfect production conditions: reducing traffic to degraded engines, limiting retries to avoid retry storms and shedding load when demand exceeds capacity. The talk presents these as practical defenses around a routing policy, rather than a guarantee that every request can be served under failure.
Watch on YouTube




