Sitanshu Gupta on Scaling AI Inference Workloads

AI Engineer15m 22s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Sitanshu Gupta lays out CoreWeave's serverless and dedicated inference models. Serverless abstracts hardware and charges by token, with a provisioned-throughput option for predictable demand; dedicated customers reserve GPU capacity and control their model deployments.

    Gupta compares agentic, chat, live voice and video, and batch workloads. Agentic systems can have long repeated inputs and rapid turns, real-time media is especially sensitive to latency, and batch jobs have flexible deadlines. Those differences shape gateway controls, routing and capacity scheduling.

    He explains KV cache locality, storage offload, optional prefill/decode separation, quantization and speculative decoding as performance choices. Reusing cached prefixes avoids costly prefills, while reserved GPUs can run batch work during quieter hours. The right configuration depends on workload shape.

    Original YouTube thumbnailWatch on YouTube