The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
FreqBLiMP is a frequency‑controlled extension of the BLiMP minimal‑pair benchmark that regenerates all 67 paradigms under explicit Zipf‑frequency regimes while preserving grammatical contrasts. The study evaluates multiple open‑weight LLM families and finds that lower lexical frequency consistently reduces sentence likelihood, yet overall contrastive acceptability accuracy drops only modestly. However, the stability in aggregate accuracy hides significant variability across linguistic phenomena, with models remaining robust on overt morphosyntactic generalization but degrading on lemma‑specific tasks.
By Tyrone White, Yuki Arase
The paper investigates how adjectival modifiers affect the semantic plausibility of events, using the Adept benchmark of 16,000 English sentence pairs that differ by a single adjective. Experiments show that sentence transformers, despite being conceptually suited to the task, underperform compared to models like RoBERTa. The authors provide an error analysis and discuss the implications of their findings for future work on balancing training and test data.
By Anna Golub, Beate Zywietz, Annerose Eichel
arXiv:2601.19926v3 Announce Type: replace-cross
Abstract: We present a systematic review of 337 articles evaluating the syntactic abilities of Transformer-based language models (TLMs), reporting on o...
By Nora Graichen, Iria de-Dios-Flores, Gemma Boleda
arXiv:2609.18284v1 Announce Type: new
Abstract: In recent years, three initiatives have emerged to develop generative language models in Hungary. The motivation behind them is the same. For Hungarian...
By M\'aty\'as Osv\'ath, Enik\H{o} H\'eja, No\'emi Ligeti-Nagy
Coordination, a fundamental linguistic structure, remains a subject of intense debate, and its exact nature continues to elude theoretical linguistics. A common view holds that only same-category constituents can be conjoined, which has been challenged by the many grammatical unlike coordinations found in natural language.
arXiv:2601. 22510v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often achieve strong benchmark accuracy yet remain brittle under small distribution shifts.
By Xingyu Zhao, Darsh Sharma, Rheeya Uppaal, Yiqiao Zhong
The paper investigates how six naturalistic and synthetic input perturbations affect decoder‑only language models at three levels: output behavior, hidden‑state geometry, and attention‑head function. Using GPT‑2 and Qwen2.5 checkpoints, the authors analyze layerwise geometry with centered kernel alignment and intrinsic dimension, and examine attention‑head responses in GPT‑2. They find that perturbation types produce distinct metric profiles that are not fully captured by output measures and vary across checkpoints, highlighting the need for multi‑level evaluation of robustness.
The paper investigates how language models avoid overgeneralizations by distinguishing between two types of indirect negative evidence: preemption and entrenchment. Through controlled rearing experiments on models trained on child‑caregiver conversations, the authors find that models do not exhibit verb‑specific preemption but show weak abstract preemption. Analysis of training dynamics suggests that competing structures act as indirect positive evidence rather than negative in the verb‑specific condition.
By Yixuan Wang, Freda Shi, Kanishka Misra
arXiv:2603.12050v3 Announce Type: replace
Abstract: Translated texts exhibit systematic differences from comparable texts originally written in the target language. Explaining this phenomenon, common...
By Maria Kunilovskaya
arXiv:2510. 22014v2 Announce Type: replace-cross Abstract: Discrete optimization-based jailbreaking attacks on large language models aim to generate short, nonsensical suffixes that, when appended onto input prompts, elicit disallowed content.
By Sarah Ball, Niki Hasrati, Alexander Robey, Avi Schwarzschild, Frauke Kreuter, Zico Kolter, Andrej Risteski
arXiv:2608. 03803v1 Announce Type: cross Abstract: Multilingual language models are deployed across a hundred or more languages, yet most benchmarks test whether a model can perform a task _in_ a language rather than whether it commands the language itself, conflating fluency with proficiency.
By Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y