arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).
By Yan Zhou
Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear...
arXiv:2603. 03612v3 Announce Type: replace Abstract: The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability.
By William Merrill, Hongjian Jiang, Yanhong Li, Anthony Lin, Ashish Sabharwal
The paper introduces a new language modeling approach that combines the benefits of Transformers and recurrent models by using balanced-tree recursive operators for autoregressive prediction. This method achieves logarithmic depth and linear runtime, allowing all prefix representations to be computed efficiently. Experiments show strong length extrapolation and performance close to ALiBi-based Transformers, suggesting it could serve as a viable alternative architecture for language modeling.
By Yiqin Wang, Nuri Cingillioglu, Charles Pert
arXiv:2606. 30461v1 Announce Type: new Abstract: State space models (SSMs) have emerged as efficient linear-time alternatives to attention for long-sequence modeling.
By Thai-Khanh Nguyen, Ngoc-Bich-Uyen Vo, Thieu N. Vo, Tan M. Nguyen, Cuong Pham