Weight Folding, CUDA Streams, and an AI Inference Bug - Filip Makraduli

AI Engineer17m 18s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Filip Makraduli presents an approach to reducing RMSNorm overhead in transformer inferenceAI inference is the process of running a trained model on new input to produce a prediction, classification, generated response or action.. Although normalization accounts for a small share of arithmetic, repeated kernel launches, memory movement and waiting can consume noticeable wall-clock time during decoding.

    The approach folds a normalization gain into a following matrix weight ahead of execution, defers a scalar division so vector and matrix work can overlap, and can remove a redundant normalization in suitable architectures. The speaker distinguishes the simple weight transformation from the lower-level GPU kernel workA graphics processing unit kernel is a function compiled to run many parallel operations on a graphics processor or related accelerator. needed for parallel execution.

    In a CUDA implementation, generated text began repeating earlier output because two streams joined without an explicit wait. Short tests did not reveal the race, while longer generation exposed a one-step lag. Waiting for both operations before post-scaling fixed the stale read.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Filip Makraduli in a blue shirt rests his hand on his chin beside the white and blue headline THE INFERENCE BUG on black. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 19 September 2026 and duration 17m 18s.

    Filip Makraduli explains how weight folding and overlapping GPU work can reduce RMSNorm overhead, then traces a stale-output bug to missing CUDA stream synchronization.