Du'an Lightfoot and Khaja Omer compare dedicated GPU inference with hosted model APIs. Sustained demand and control over data can favor a managed GPU environment, while intermittent traffic, stronger hosted models or limited operational expertise can favor an API. Their workshop uses isolated Kubernetes namespaces and an OpenAI-compatible vLLM endpoint.
The presenters measure time to first token, decode latency and aggregate throughput rather than treating tokens per second as a single performance score. Model weights and growing KV caches compete for GPU memory, while prompts, tool calls and concurrent users change the available capacity. They discuss batching, prefix caching and observability as parts of that workload budget.
A redeployment from BF16 to a prequantized FP8 model improves performance in their test. Omer emphasizes checking the actual hardware and evaluating the application's structured outputs and tool calls, because lower precision is not a universal quality-preserving shortcut. Their speculative-decoding example then performs worse, demonstrating why draft compatibility and token acceptance must be measured instead of assuming a speedup.
The final section tunes vLLM around a throughput-latency knee, changing configuration and rerunning load tests. A notebook failure changes the environment, and the final comparison uses BF16 rather than the earlier FP8 model. The recording shows inference configuration and tuning, not the complete agent deployment originally planned for the workshop.
Watch on YouTube




