The study evaluates synthetic pre‑pretraining (PPT) across a wide range of models (500 M–7 B parameters) and training budgets (up to 100 B tokens). Results show that PPT consistently improves downstream performance and token efficiency, saving at least 21 B tokens at the 3 B scale, but these gains do not appear to stem from a grammatical prior. Instead, PPT benefits arise from tasks that enhance long‑range retrieval, and the improvements remain robust across diverse data mixtures, diminishing only when web text is omitted.
By Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal \v{S}tef\'anik, Aline Villavicencio, Nikolaos Aletras
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
By Haris Riaz, Hyungji Kim, Mihai Surdeanu
arXiv:2606. 25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention.
By Shen Nie, Qiyang Min, Shaoxuan Xu, Zihao Huang, Yuxuan Song, Yong Shan, Yankai Lin, Wayne Xin Zhao, Chongxuan Li, Ji-Rong Wen
arXiv:2608.29034v1 Announce Type: cross
Abstract: A wide range of methods have been proposed for interpreting language models, delivering important insights into their inner workings. However, differ...
By Zhang Enyan, R. Thomas McCoy
arXiv:2602. 06065v3 Announce Type: replace-cross Abstract: Understanding how the structure of language can be learned from sentences alone is a central question in both cognitive science and machine learning.
By Jack T. Parley, Francesco Cagnetta, Matthieu Wyart
The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.
By Yixuan Wang, Freda Shi, Kanishka Misra
arXiv:2601. 22594v2 Announce Type: replace-cross Abstract: The high-level concepts that a neural network uses to perform computation need not be aligned to individual neurons (Smolensky, 1986).
By Aryaman Arora, Zhengxuan Wu, Jacob Steinhardt, Sarah Schwettmann
arXiv:2606.18389v2 Announce Type: replace
Abstract: Large language models (LLMs) have become an effective tool for synthetic data generation, including for low-resource languages, where generated dat...
By Jan Cegin, Daniil Gurgurov, Yusser Al Ghussin, Simon Ostermann
The paper introduces a probe‑free method called the Neuron Separability Index (NSI) to assess how individual neurons in large language models distinguish grammatical from ungrammatical sentences using linguistic minimal pairs. Across 68 linguistic paradigms and seven model checkpoints, the study finds that while raw separability for morphology and syntax peaks early, single‑unit selectivity is sparse and weak, with rare strongly selective "grandmother neurons." Moreover, the research shows a dissociation between whole‑vector linear separability, single‑neuron selectivity, and behavioral competence, and demonstrates that targeted ablations can further separate activation selectivity from causal reliance.
The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
The paper introduces a probe‑free method called the Neuron Separability Index (NSI) to assess how individual neurons in Large Language Models (LLMs) distinguish grammatical from ungrammatical constructions. Using linguistic minimal pairs across 68 paradigms and seven checkpoints, the study finds that raw separability peaks earlier for morphological and syntactic distinctions, but after permutation normalization, single‑unit selectivity is sparse, weak, and narrowly tuned, with rare strongly selective "grandmother neurons". Additionally, whole‑vector linear separability, single‑neuron selectivity, and behavioral competence are largely dissociated, and targeted ablations further separate activation selectivity from causal reliance.
By Linyang He, Nima Mesgarani
The study explores how neural agents develop dependency length minimization (DLM) in artificial languages using a recurrent neural network framework. By manipulating processing constraints such as listening noise, speaker capacity, and incremental sentence processing, the researchers find that DLM emerges only under incremental processing pressure, while other factors produce varied word‑order preferences. These results suggest that human cognitive processing limits may influence the emergence of DLM in language.
By Yuqing Zhang, Tessa Verhoef, Gertjan van Noord, Arianna Bisazza