arXiv Machine Learning By George Fountzoulas

Kathleen Writes: Autoregressive Generation and Data Scaling Without Attention

Read the original on arXiv Machine Learning →

arXiv:2608. 04678v1 Announce Type: cross Abstract: Papers 1-2 of the Kathleen series showed that a byte-level, attention-free architecture built from a wavetable encoder and multi-scale reverberant state can match strong baselines on classification at ~450-700K parameters, without pretraining.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Relational Attention for Data-Efficient Language Modeling

Relational BabyLM is a decoder‑only Transformer that replaces standard self‑attention with a Dual Attention Transformer (DAT) to separate object‑level lexical features from structural/relational information. The model incorporates a Next‑Latent Prediction objective to compress history into a dense belief state and introduces a RoPE‑based symbol‑retrieval mechanism. On the BabyLM 2026 challenge, the best model ranks 6th overall and 3rd on the NLP‑task subset, outperforming GPT‑2 on most benchmarks and achieving the highest EWoK score among strict‑track entries.

By Adrian Brasoveanu, Ece Takmaz, Jakub Dotla\v{c}il
arXiv Computation and Language
Sep 4

Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding

The paper introduces Variance‑Calibrated Modulation (VCM), a training‑free pre‑decoding technique that reshapes language model probability distributions before truncation. VCM uses two dynamic mechanisms: a Contextual Searchlight via PMI to suppress stopwords and highlight context‑relevant tokens, and an Adaptive Self‑Debiasing that applies scale‑invariant penalization based on real‑time logit standard deviation. Experiments on open‑ended generation, factual QA, and mathematical reasoning show that VCM consistently reduces the likelihood trap, improving diversity, coherence, and reasoning accuracy with minimal computational cost.

By Yuanhao Ding, Meimingwei Li, Esteban Garces Arias, Matthias A{\ss}enmacher, Christian Heumann, Chongsheng Zhang
arXiv Machine Learning
Sep 23

Latest Exact Match Attention

arXiv:2609.25802v1 Announce Type: new Abstract: We introduce latest exact match attention (LEMA), an attention variant for transformers where queries and keys are binarized and each query attends onl...

By Moritz Br\"osamle
arXiv Computation and Language
Sep 16

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.

By Eduardo Novaes Hering