arXiv AI By R\'ois\'in Luo

ZetaGPT: A Reference Implementation of Positional--Encoding--Free State--Space--Attention Language Models

Read the original on arXiv AI →

arXiv:2608. 09432v1 Announce Type: cross Abstract: Transformer-based language models rely on self-attention, whose computation is permutation-equivariant and therefore lacks an intrinsic mechanism for representing token order.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 8

RePo: Language Models with Context Re-Positioning

arXiv:2512. 14391v3 Announce Type: replace-cross Abstract: In-context learning is fundamental to modern Large Language Models (LLMs); however, prevailing architectures impose a rigid and fixed contextual structure by assigning linear or constant positional indices.

By Huayang Li, Tianyu Zhao, Deng Cai, Richard Sproat
arXiv Machine Learning
Sep 18

Relational Attention for Data-Efficient Language Modeling

Relational BabyLM is a decoder‑only Transformer that replaces standard self‑attention with a Dual Attention Transformer (DAT) to separate object‑level lexical features from structural/relational information. The model incorporates a Next‑Latent Prediction objective to compress history into a dense belief state and introduces a RoPE‑based symbol‑retrieval mechanism. On the BabyLM 2026 challenge, the best model ranks 6th overall and 3rd on the NLP‑task subset, outperforming GPT‑2 on most benchmarks and achieving the highest EWoK score among strict‑track entries.

By Adrian Brasoveanu, Ece Takmaz, Jakub Dotla\v{c}il
Hugging Face Trending Papers
5d ago

Not All Thinking is Created Equal: Latent Reasoning Discovers a Recurrent Search Algorithm for Depth Generalization

The paper investigates whether different forms of intermediate computation in large language models—such as token-based traces, pause tokens, and latent reasoning—rely on the same underlying mechanism. By training five variants of GPTNeoX on an extended multi-hop reasoning task, the authors find that while vanilla, Chain-of-Thought, and Pause Token models perform well on in-distribution data, they fail to generalize to longer-hop out-of-distribution problems. In contrast, latent-reasoning models exhibit better depth generalization, with causal analysis revealing a sparse recurrent search circuit that implements forward reachability propagation across the graph.