DeepSeek's New Architectures: mHC, Engram and Conditional Memory

The Pretrained Pod1h
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Richard Diehl Martinez and Pierce Freeman begin with residual connections: adding a layer's input to its transformed output. They use canal and telephone analogies to explain why identity paths can help information and gradients move through deep networks.

    The first paper, mHC, addresses the instability introduced by richer hyper-connections. Its constrained stream-mixing design aims to preserve useful residual properties while supporting training at scale. The hosts place the research within a broader debate about compute efficiency and model quality.

    They connect sparse architectures to the search for smaller models that retain useful performance. Their discussion of the lottery-ticket hypothesis and national research incentives is exploratory, not proof that any large model has an easily extracted tiny equivalent.

    The Engram paper introduces conditional memory through learned lookup embeddings. The hosts compare this idea with retrieving prepared ingredients rather than rebuilding everything from scratch, then contrast it with attention's growing cache and conventional document-based retrieval.

    The final discussion asks where memory modules might help personalized or enterprise systems, including auditability and access control. These are proposed applications, not demonstrated Engram features. The hosts remain uncertain about product fit and the operational complexity that additional memory can introduce.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Pierce Freeman and Richard Diehl Martinez in blue and white tops against black, one touching his chin, alongside the blue and white headline "DEEPSEEK'S NEW BLUEPRINT". Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 30 January 2026 and duration 1h.

    The hosts explore two DeepSeek papers: mHC for more stable hyper-connections and Engram for learned lookup memory. They discuss how these approaches differ from conventional residual connections, attention caches and document retrieval.