arXiv:2607. 08152v1 Announce Type: cross Abstract: On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance.
By Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Tien Nguyen, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv:2608.30583v1 Announce Type: new
Abstract: Standard language proficiency tests rely on linguistic tasks such as vocabulary, grammar and reading comprehension quizzes. An alternative, cognitively...
By Shachar Frenkel, Ido Falah, Omer Shubi, Yevgeni Berzak
arXiv:2505.12196v2 Announce Type: replace
Abstract: The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conject...
By Yi-Chien Lin, Hongao Zhu, William Schuler
The study introduces a latent‑process multinomial processing tree (MPT) model to analyze human reading and comprehension of garden‑path sentences across four reading paradigms (eye tracking, uni‑ and bidirectional self‑paced reading, Maze). The model separates the likelihood of an incorrect initial analysis, the cost of encountering an incompatible continuation, and the cost of syntactic reanalysis, yielding more realistic parameter estimates when inattentive trials are considered. Cross‑validation shows that this MPT model predicts human reading patterns and end‑of‑trial judgments better than a model relying solely on large‑language‑model (LLM) surprisal, and that incorporating surprisal as an additional predictor further improves fit.
By Dario Paape, Tal Linzen, Shravan Vasishth
arXiv:2609.18011v1 Announce Type: new
Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
By Nan Li, Albert Gatt, Massimo Poesio
arXiv:2608.22452v1 Announce Type: new
Abstract: Surprisal, the negative log-probability a language model assigns to a word given its preceding context, reliably predicts adult reading times. Does it...
By Francisco Portillo L\'opez
Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, while exact, full-vocabulary Kullback--Leibler (KL) training dominates memory.
arXiv:2607.08152v2 Announce Type: replace-cross
Abstract: Predicting comprehension from eye movements could support adaptive reading interfaces. We present LEXIC, a compact recurrent model that predi...
By Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Hyunsu Go, Eunseob Choi, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv:2608. 10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network.
By Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster
The paper introduces Sparse Readout Prism (SRP), a method that decomposes a language model’s readout matrix into sparse features, allowing logit‑lens scores to be expressed as sums of feature contributions. SRP reveals that lens readings depend on the corpus used to fit the readout, a phenomenon called corpus conditionality, and that the dominant readout feature remains stable across different corpora. By replacing the original readout with SRP’s sparse approximation, the authors recover 8.9–17.3 percentage points more of the tested logit differences than six geometric‑relation baselines, and ablating features shifts logit differences proportionally to their SRP contributions.
By Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane
The paper investigates how general vision‑language models (VLMs) develop specialized optical character recognition (OCR) capabilities. By applying a causal intervention protocol, the authors identify sparse, stable OCR‑head sets in several VLMs and show that these heads largely overlap with textual retrieval/copy heads found in general VLMs. The study concludes that full‑sequence OCR functions as a dense multimodal copy‑and‑paste mechanism, and that when a VLM is fine‑tuned for OCR, it largely preserves the same head identities while redistributing their functional and causal strengths.
By Yuanxiang Huangfu, Hanmeng Zhong, Linqing Chen, Jeffrey Tiong Jee Hui
The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.
By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li