The study introduces a latent‑process multinomial processing tree (MPT) model to analyze human reading and comprehension of garden‑path sentences across four reading paradigms (eye tracking, uni‑ and bidirectional self‑paced reading, Maze). The model separates the likelihood of an incorrect initial analysis, the cost of encountering an incompatible continuation, and the cost of syntactic reanalysis, yielding more realistic parameter estimates when inattentive trials are considered. Cross‑validation shows that this MPT model predicts human reading patterns and end‑of‑trial judgments better than a model relying solely on large‑language‑model (LLM) surprisal, and that incorporating surprisal as an additional predictor further improves fit.
By Dario Paape, Tal Linzen, Shravan Vasishth
The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.
By Andr\'es Bux\'o-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard, Cassandra L. Jacobs
arXiv:2608. 11138v1 Announce Type: cross Abstract: We propose that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is \emph{fragile} under perturbation of its attention pathways.
By Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon
arXiv:2606.27206v2 Announce Type: replace
Abstract: Garden path sentences present a processing difficulty for humans--- the sentence prefix leads the listener towards one interpretation, until the li...
By Alan Zhou, Milo\v{s} Stanojevi\'c, John T. Hale
arXiv:2608.22452v1 Announce Type: new
Abstract: Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it...
By Francisco Portillo L\'opez
arXiv:2509. 14704v3 Announce Type: replace Abstract: Benchmark saturation and training-data contamination increasingly obscure whether reported gains in large language models (LLMs) reflect genuine advances in reasoning or familiarity with recurring patterns in benchmark problems.
By Masaharu Mizumoto, Dat Nguyen, Zhiheng Han, Xingfu Li, Yo Nakawake, Le Minh Nguyen