arXiv AI

When transformers learn "impossible" languages, what do they learn?

arXiv:2606. 30815v1 Announce Type: cross Abstract: Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans.

arXiv Computation and Language
Aug 31

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.

By Titus von der Malsburg, Sebastian Pad\'o
arXiv AI
Sep 10

FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

FreqBLiMP is a frequency‑controlled extension of the BLiMP minimal‑pair benchmark that regenerates all 67 paradigms under explicit Zipf‑frequency regimes while preserving grammatical contrasts. The study evaluates multiple open‑weight LLM families and finds that lower lexical frequency consistently reduces sentence likelihood, yet overall contrastive acceptability accuracy drops only modestly. However, the stability in aggregate accuracy hides significant variability across linguistic phenomena, with models remaining robust on overt morphosyntactic generalization but degrading on lemma‑specific tasks.

By Tyrone White, Yuki Arase
arXiv Computation and Language
Sep 17

Modelling Adjectival Modification Effects on Semantic Plausibility

The paper investigates how adjectival modifiers affect the semantic plausibility of events, using the Adept benchmark of 16,000 English sentence pairs that differ by a single adjective. Experiments show that sentence transformers, despite being conceptually suited to the task, underperform compared to models like RoBERTa. The authors provide an error analysis and discuss the implications of their findings for future work on balancing training and test data.

By Anna Golub, Beate Zywietz, Annerose Eichel
Hugging Face Trending Papers
Jul 22

Exposure is Optional: Learning Unlike Coordination in Language Models

Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.

Hugging Face Trending Papers
Sep 3

How Perturbations Propagate: A Multi-Level Analysis of Robustness in Large Language Models

The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.

arXiv Computation and Language
Sep 3

Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization

The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.

By Yixuan Wang, Freda Shi, Kanishka Misra
arXiv Machine Learning
Aug 5

M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models

arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.

By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y