arXiv:2607. 17419v1 Announce Type: cross Abstract: Linear attention promises constant-time recurrent inference but degrades sharply on associative recall.
By Ayoub Ghriss, Sourav Chakraborty
arXiv:2608.30386v1 Announce Type: cross
Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substa...
By Yanqi Yu, Pingwei Sun, Jianchao Tan, Tao Zhang, Yuchen Xie, Xunliang Cai, Yao Liu
Fast Weight Attention for Continual Learning introduces recurrent fast‑weight memories and selective state‑space models that compress expanding context into a fixed‑size recurrent state, enabling an online learning rule for state transitions. The paper derives normalized first‑order updates for squared‑error regression and negative inner‑product objectives, presenting several variants (Falcon‑1, Falcon‑2, Falcon‑3 and their inner‑product counterparts) with recurrent, masked‑parallel, and chunk‑parallel implementations. These methods demonstrate competitive performance in language modeling and improved length extrapolation on variable‑digit addition tasks.
By Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active overwriting of stale ones. In our diagnostic suite, the strongest efficient baselines tend to solve only one side well.
arXiv:2607. 21000v1 Announce Type: new Abstract: Long-sequence memory tracking places two opposing demands on a recurrent state: near-lossless retention of stored bindings over long horizons, and active overwriting of stale ones.
By Hyuk Lim, Seunghyun Yoon
arXiv:2606. 09862v1 Announce Type: cross Abstract: The Softmax Attention operation in Transformer language models has a quadratic complexity in the sequence length and a growing state size in the form of KV cache, which becomes a bottleneck in long context scenarios.
By Axel Laborieux, Christos Sourmpis, Juan Gabriel Kostelec, Qinghai Guo