Homa: The End of TCP for AI Clusters - John Ousterhout, Stanford

AI Engineer18m 48s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    John Ousterhout argues that inferenceAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action. and agentic workloads increasingly depend on short coordination messages, even while training still moves large amounts of data. When distributed GPU jobs must synchronize between compute phases, a slow message can leave processors idle and limit overall throughput.

    The talk traces tail latency to incast, where large transfers fill a switch queue and short messages wait behind them. It also explains why sender-side congestion control reacts after network feedback and why byte streams make it difficult to prioritize individual short messages.

    John Ousterhout presents Homa as a message-based transport that knows message lengths, lets receivers pace additional packets with grants, and uses switch priority queues to favor short exchanges. In one benchmark described in the talk, Homa's 99th-percentile latency for short messages is reported at less than 100 microseconds, versus more than one millisecond for TCP, roughly a 13-fold difference. That is the speaker's sample benchmark, not an independently verified general performance guarantee.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    John Ousterhout in a dark shirt on a black background beside the blue-and-white headline HOMA FOR AI CLUSTERS. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 17 September 2026 and duration 18m 48s.

    John Ousterhout argues that Homa can cut the tail latency of short AI-cluster messages by replacing sender-led congestion control with message-aware receiver scheduling.