arXiv Machine Learning
Sep 24

Log-Depth Recurrent Language Modeling

The paper introduces a new language modeling approach that combines the benefits of Transformers and recurrent models by using balanced-tree recursive operators for autoregressive prediction. This method achieves logarithmic depth and linear runtime, allowing all prefix representations to be computed efficiently. Experiments show strong length extrapolation and performance close to ALiBi-based Transformers, suggesting it could serve as a viable alternative architecture for language modeling.

By Yiqin Wang, Nuri Cingillioglu, Charles Pert
Hugging Face Trending Papers
Jun 16

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers. In language modeling, \textbf{transformers} have emerged as the dominant architecture, with early layers capturing local syntactic patterns and later layers encoding more complex clause-level dependencies.

arXiv Machine Learning
Jun 17

An expressivity analysis of hierarchical modelling in deep transformers via bounded-depth grammars

arXiv:2606. 17522v1 Announce Type: cross Abstract: Deep neural networks are widely believed to derive their expressive power from their ability to form \textbf{hierarchical representations}, capturing progressively more abstract and compositional features across layers.

By Vinoth Nandakumar, Qiang Qu, Pramod Thebe, Sakshi Khachariya, Tongliang Liu
arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi