arXiv Machine Learning

Relational Attention for Data-Efficient Language Modeling

Relational BabyLM is a decoder‑only Transformer that replaces standard self‑attention with a Dual Attention Transformer (DAT) to separate object‑level lexical features from structural/relational information. The model incorporates a Next‑Latent Prediction objective to compress history into a dense belief state and introduces a RoPE‑based symbol‑retrieval mechanism. On the BabyLM 2026 challenge, the best model ranks 6th overall and 3rd on the NLP‑task subset, outperforming GPT‑2 on most benchmarks and achieving the highest EWoK score among strict‑track entries.

arXiv Machine Learning
Sep 23

Latest Exact Match Attention

arXiv:2609.25802v1 Announce Type: new Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...

By Moritz Br\"osamle
arXiv Computation and Language
Sep 16

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.

By Eduardo Novaes Hering
arXiv Machine Learning
Jun 2

Resonant Context Anchoring: Decoupling Attention Routing and Signal Gain at Inference Time

arXiv:2606. 01923v1 Announce Type: cross Abstract: Large Language Models (LLMs) frequently exhibit "contextual disregard" when faced with input evidence that conflicts with their internal parametric memory, leading to persistent factual hallucinations.

By Mingkuan Zhao, Yide Gao, Wentao Hu, Suquan Chen, Tianchen Huang, Zhenhua An, Zetao Chang, Xiayu Sun, Yuheng Min