arXiv AI

FreqBLiMP: Frequency-Controlled Minimal Pairs Reveal Robustness and Fragility of LLMs Under Lexical Rarity

FreqBLiMP is a frequency‑controlled extension of the BLiMP minimal‑pair benchmark that regenerates all 67 paradigms under explicit Zipf‑frequency regimes while preserving grammatical contrasts. The study evaluates multiple open‑weight LLM families and finds that lower lexical frequency consistently reduces sentence likelihood, yet overall contrastive acceptability accuracy drops only modestly. However, the stability in aggregate accuracy hides significant variability across linguistic phenomena, with models remaining robust on overt morphosyntactic generalization but degrading on lemma‑specific tasks.

arXiv Computation and Language
Aug 31

Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects

The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.

By Titus von der Malsburg, Sebastian Pad\'o
arXiv Computation and Language
5d ago

How Much Does Corpus Choice Change Dependency-Distance Estimates?

The study examined how the choice of corpus affects estimates of dependency distance in language. By comparing 38 pairs of treebanks from the same language, the authors found that cross-treebank agreement was only moderate, with nearly 40% of language orderings reversed when switching treebanks. Treebank selection explained about 29% of the variance, a discrepancy that far exceeds within-treebank sampling error and persists across multiple preprocessing settings, yet all treebanks still supported the principle of dependency-length minimization.

By Sirui Chen
arXiv AI
Jul 9

DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation

arXiv:2607. 07669v1 Announce Type: cross Abstract: Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed.

By Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia, Aditya Joshi, Lu Yin
arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv AI
Sep 3

CroCo: Cross-Lingual Contrastive Preference Tuning on Self-Generations

CroCo introduces cross‑lingual contrastive preference tuning on self‑generations, extending prior English‑only methods to 14 high‑ and low‑resource languages. A reward model trained solely on English preferences, applied to a multilingual base, yields effective within‑language rankings and improves performance in both monolingual and multilingual settings without catastrophic forgetting. The approach requires on‑policy data; off‑policy responses and online preference optimization offer limited gains, yet on structured tasks CroCo matches or surpasses the base model in most languages, and on open‑ended generation it wins 28/30 judge evaluations across 15 languages.

By Mike Zhang, Ali Basirat, Desmond Elliott