Agents That Own Their Inference - Du'an Lightfoot & Khaja Omer, Akamai Technologies

AI Engineer1h 50m
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Du'an Lightfoot and Khaja Omer compare dedicated GPU inference with hosted model APIs. Sustained demand and control over data can favor a managed GPU environment, while intermittent traffic, stronger hosted models or limited operational expertise can favor an API. Their workshop uses isolated Kubernetes namespaces and an OpenAI-compatible vLLM endpoint.

    The presenters measure time to first token, decode latency and aggregate throughput rather than treating tokens per second as a single performance score. Model weights and growing KV caches compete for GPU memory, while prompts, tool calls and concurrent users change the available capacity. They discuss batching, prefix caching and observability as parts of that workload budget.

    A redeployment from BF16 to a prequantized FP8 model improves performance in their test. Omer emphasizes checking the actual hardware and evaluating the application's structured outputs and tool calls, because lower precision is not a universal quality-preserving shortcut. Their speculative-decoding example then performs worse, demonstrating why draft compatibility and token acceptance must be measured instead of assuming a speedup.

    The final section tunes vLLM around a throughput-latency knee, changing configuration and rerunning load tests. A notebook failure changes the environment, and the final comparison uses BF16 rather than the earlier FP8 model. The recording shows inference configuration and tuning, not the complete agent deployment originally planned for the workshop.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Portraits of Du'an Lightfoot, Khaja Omer against black with the headline OWN YOUR INFERENCE in attention blue and white. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 11 October 2026 and duration 1h 50m.

    Du'an Lightfoot and Khaja Omer show how memory budgets, workload evaluations and latency-throughput measurements guide self-hosted inference, including optimizations that fail in practice.