arXiv Machine Learning By Gyuryang Heo, Timothy Ngotiaoco, Kazuki Irie, Samuel J. Gershman, Bernardo L. Sabatini

Why Depth Matters in Parallelizable Sequence Models: A Lie Algebraic View

Read the original on arXiv Machine Learning →

arXiv:2603. 05573v2 Announce Type: replace Abstract: Scalable sequence models, such as Transformer variants and structured state-space models, often trade expressivity power for sequence-level parallelism, which enables efficient training.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 9

MinMax Recurrent Neural Cascades

arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.

By Alessandro Ronca
arXiv AI
Aug 7

The Impossibility Triangle of Long-Context Modeling

arXiv:2605. 05066v2 Announce Type: replace-cross Abstract: We identify and prove a fundamental trade-off governing long-sequence models: no model can simultaneously achieve (i) per-step computation independent of sequence length (Efficiency), (ii) state size independent of sequence length (Compactness), and (iii) the ability to recall a number of historical facts proportional to sequence length (Recall).

By Yan Zhou
Hugging Face Trending Papers
Sep 23

Log-Depth Recurrent Language Modeling

Language modeling using Transformers has become commonplace despite their fixed computational depth and quadratic runtime with respect to input tokens. Recurrent models on the other hand offer linear...

arXiv Machine Learning
Jun 3

Why Are Linear RNNs More Parallelizable?

arXiv:2603. 03612v3 Announce Type: replace Abstract: The community is increasingly exploring linear RNNs (LRNNs) as language models, motivated by their expressive power and parallelizability.

By William Merrill, Hongjian Jiang, Yanhong Li, Anthony Lin, Ashish Sabharwal
arXiv Machine Learning
Sep 24

Log-Depth Recurrent Language Modeling

The paper introduces a new language modeling approach that combines the benefits of Transformers and recurrent models by using balanced-tree recursive operators for autoregressive prediction. This method achieves logarithmic depth and linear runtime, allowing all prefix representations to be computed efficiently. Experiments show strong length extrapolation and performance close to ALiBi-based Transformers, suggesting it could serve as a viable alternative architecture for language modeling.

By Yiqin Wang, Nuri Cingillioglu, Charles Pert