Linear attention uses mathematical rearrangements, kernel functions or recurrent summaries to avoid constructing the complete token-by-token attention matrix. This can make long sequences cheaper to process in memory and computation.
The efficiency gain comes with design tradeoffs. Approximations may lose some expressive detail, and different linear attention methods behave differently on recall and reasoning tasks. Hybrid architectures often combine linear and conventional or sparse attention to balance quality and cost.
