Károly Zsolnai-Fehér Explains DeepSeek's Shared KV Memory

Two Minute Papers5m 10s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Károly Zsolnai-Fehér describes DeepSeek 4.1 Flash's reported speed, selected benchmark results and native visual understanding. He emphasizes that comparisons vary by test rather than claiming the model wins universally.

    Károly Zsolnai-Fehér explains the encoder-decoder idea behind CSA2: an encoder creates shared global memory that the decoder reads, reducing the need for every layer to maintain its own history. He presents the smaller KV cache as a significant memory-efficiency improvement, not a claim that the full model is small enough for ordinary home hardware.

    Károly Zsolnai-Fehér discusses image-to-game experiments and compares other models on reproducing a physics paper. He also cautions that DeepSeek can consume many reasoning tokens and still has a very large parameter footprint, so lower memory overhead does not eliminate deployment costs.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Károly Zsolnai-Fehér in a photographic portrait composition on black beside the blue-and-white headline “SHARED AI MEMORY”. Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 18 September 2026 and duration 5m 10s.

    Károly Zsolnai-Fehér explains how shared memory across layers can shrink DeepSeek's KV cache, while a huge parameter count and heavy reasoning still limit local use.