John Ousterhout argues that inferenceAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action. and agentic workloads increasingly depend on short coordination messages, even while training still moves large amounts of data. When distributed GPU jobs must synchronize between compute phases, a slow message can leave processors idle and limit overall throughput.
The talk traces tail latency to incast, where large transfers fill a switch queue and short messages wait behind them. It also explains why sender-side congestion control reacts after network feedback and why byte streams make it difficult to prioritize individual short messages.
John Ousterhout presents Homa as a message-based transport that knows message lengths, lets receivers pace additional packets with grants, and uses switch priority queues to favor short exchanges. In one benchmark described in the talk, Homa's 99th-percentile latency for short messages is reported at less than 100 microseconds, versus more than one millisecond for TCP, roughly a 13-fold difference. That is the speaker's sample benchmark, not an independently verified general performance guarantee.
Watch on YouTube




