Linear Attention Architectures: Mechanisms, Trade-offs, and Cross-Layer Routing
paper Your tags
Your notes
The reference controlled comparison for the DeltaNet-family attention wave: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2 vs softmax attention in a unified recurrent-memory notation, at 350M params / 15B tokens (plus 1.3B/3B DeltaNet runs), with cross-layer routing experiments. KDA + Muon reaches the lowest validation loss in scope. Imanol Schlag's group, ETH Zurich / ETH AI Center.