Qwen's Leadership Changes and Hybrid Attention Explained

The Pretrained Pod10m 47s
0 comments · 0 votesOpen discussionClose discussion
Sign in to join the discussion

    Video summary

    Pierce Freeman and Richard Diehl Martinez discuss reports that Junyang Lin and other Qwen researchers were leaving Alibaba. They acknowledge uncertainty about the situation and possible reorganization, contrasting the news with enthusiasm for the recently released Qwen3.5 family.

    The episode explains why full attention becomes expensive as input length grows and why recurrent, linear-attention approaches can reduce some of that cost. The hosts use copying and associative recall as examples of information that compressed memory can struggle to retain.

    Gated DeltaNet provides a way to update and gate a running memory state. Qwen3.5 combines that approach with full attention rather than relying on either alone, seeking a practical balance between efficiency and retrieval. These are architectural trade-offs, not proof that every linear-attention model is incapable of copying.

    The hosts end with hopes for continued open-model development and competing interpretations of a reported follow-up message. Their expectations about future leadership, releases and licensing remain unresolved at the time of the episode; the discussion does not establish that the entire Qwen team quit.

    Original YouTube thumbnailWatch on YouTube

    Share this page

    Pierce Freeman, Richard Diehl Martinez, Junyang Lin against black with the blue and white headline "QWEN'S TEAM, HYBRID MODEL BET". Framed in blue with WWW.ARTIFICIAL-INTELLIGENCE.VIDEO, 13 March 2026 and duration 10m 47s.

    Reports of Qwen leadership changes frame a technical discussion of linear attention, associative recall and Gated DeltaNet. The hosts are enthusiastic about efficient open models but uncertain about the team's future.