The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
By Titus von der Malsburg, Sebastian Pad\'o
The paper investigates energy-based transformers as predictors of reading difficulty, extending the use of transformer language models in psycholinguistics. It demonstrates that the energy measure from these models robustly predicts reading times across multiple corpora, outperforming traditional metrics like surprisal and attention entropy. In a controlled experiment on relative clause processing, energy captures known asymmetries, suggesting it may unify previously complementary predictors.
By Jakub Dotlacil, Ece Takmaz
arXiv:2505.12196v2 Announce Type: replace
Abstract: The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conject...
By Yi-Chien Lin, Hongao Zhu, William Schuler
arXiv:2609.36214v1 Announce Type: new
Abstract: Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematic...
By Yusheng Zhou, Eleanor Lin, David Jurgens
The study introduces a latent‑process multinomial processing tree (MPT) model to analyze human reading and comprehension of garden‑path sentences across four reading paradigms (eye tracking, uni‑ and bidirectional self‑paced reading, Maze). The model separates the likelihood of an incorrect initial analysis, the cost of encountering an incompatible continuation, and the cost of syntactic reanalysis, yielding more realistic parameter estimates when inattentive trials are considered. Cross‑validation shows that this MPT model predicts human reading patterns and end‑of‑trial judgments better than a model relying solely on large‑language‑model (LLM) surprisal, and that incorporating surprisal as an additional predictor further improves fit.
By Dario Paape, Tal Linzen, Shravan Vasishth
The paper introduces PILL, a new infilling technique for diffusion language models that eliminates the need for a preset initial length and reduces inference overhead. PILL uses probing-based length-free decoding, cutting down on extra forward passes and speeding up generation. Experiments across five diffusion models and eight benchmarks show PILL outperforms the strongest baseline with higher pass rates and BLEU-2 scores while running 1.82× faster.
By Haobo Xu, Sirui Chen, Yuanchen Bei, Lingjie Chen, Yuchen Yan, Dongqi Fu, Jingrui He, Hanghang Tong