DeepSeek did it again...

Matthew Berman16m 39s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    DeepSeek presents V4.1 Flash as a 552-billion-parameter mixture-of-experts model that activates only 8 billion parameters for input and 16 billion for output. Matthew Berman highlights the resulting speed and memory savings, including claims that its KV cache needs one-quarter of the HBM and that the model needs one-eighth of the SSD storage described in the comparison.

    Published benchmark tables presented in the video place DeepSeek V4.1 Flash near or above Claude Opus 5 and GPT-5.6 Sol on some suites, while showing weaker performance on Exploit Gym. These are benchmark claims discussed in the video, not independent verification by Matthew Berman.

    Matthew Berman argues that open weights, very low token pricing and reduced memory demands make the model attractive as a high-volume workhorse and a future local option after quantization. He also estimates roughly 200 output tokens per second in a simple essay test.

    The hands-on coding demonstrations are more mixed. DeepSeek V4.1 Flash produced broken Rubik's Cube simulations, created a stylized but low-detail paint result and built a configurable 3D water-impact app whose central simulation remained unconvincing. Matthew Berman concludes that the model is fast and economical, but not consistently dependable for demanding coding tasks.

    Original YouTube thumbnailWatch on YouTube