arXiv Computation and Language

Reader Proficiency Shapes Layer-wise Surprisal Profiles

The study examines how reader proficiency influences the relationship between layer-wise surprisal from large language models (LLMs) and eye-tracking gaze measures. Using the MECO L2 corpus, researchers compared high- and low-proficiency readers on first-pass gaze duration (FPGD) and total gaze duration (TGD), finding that lower-proficiency readers exhibit deeper Predictive Depth for FPGD, while TGD shows deeper Predictive Depth across both groups. The results suggest that the distribution of predictive power across LLM layers relates to the timing and breadth of reading processes and varies with reader proficiency.

arXiv AI
Jul 10

LEXIC: Lightweight Eye-tracking eXtension via Injected Complexity

arXiv:2607. 08152v1 Announce Type: cross Abstract: On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance.

By Sumin Lee, Kyeonghun Kim, Subeen Lee, Jiwon Yang, Tien Nguyen, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv Computation and Language
Sep 25

LLM surprisal is necessary but not sufficient to capture English garden-path effects: Evidence from joint latent modeling of reading paradigms

The study introduces a latent‑process multinomial processing tree (MPT) model to analyze human reading and comprehension of garden‑path sentences across four reading paradigms (eye tracking, uni‑ and bidirectional self‑paced reading, Maze). The model separates the likelihood of an incorrect initial analysis, the cost of encountering an incompatible continuation, and the cost of syntactic reanalysis, yielding more realistic parameter estimates when inattentive trials are considered. Cross‑validation shows that this MPT model predicts human reading patterns and end‑of‑trial judgments better than a model relying solely on large‑language‑model (LLM) surprisal, and that incorporating surprisal as an additional predictor further improves fit.

By Dario Paape, Tal Linzen, Shravan Vasishth
Hugging Face Trending Papers
Aug 10

Interpreting Language Model Hidden States at Scale

Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network. Trained lenses remain expensive: affine-translator parameters grow quadratically with model width, while exact, full-vocabulary Kullback--Leibler (KL) training dominates memory.

arXiv AI
Aug 12

Interpreting Language Model Hidden States at Scale

arXiv:2608. 10260v1 Announce Type: new Abstract: Lens methods interpret large language models (LLMs) by mapping intermediate activations to the output vocabulary, revealing how next-token predictions develop through the network.

By Jordan Pettyjohn, Mansi Sakarvadia, Nathaniel Hudson, Daniel McKenzie, Kyle Chard, Ian Foster
arXiv AI
Sep 3

Sparse Readout Prism: Explaining Logit-Lens Scores in Features Instead of Tokens

The paper introduces Sparse Readout Prism (SRP), a method that decomposes a language model’s readout matrix into sparse features, allowing logit‑lens scores to be expressed as sums of feature contributions. SRP reveals that lens readings depend on the corpus used to fit the readout, a phenomenon called corpus conditionality, and that the dominant readout feature remains stable across different corpora. By replacing the original readout with SRP’s sparse approximation, the authors recover 8.9–17.3 percentage points more of the tested logit differences than six geometric‑relation baselines, and ablating features shifts logit differences proportionally to their SRP contributions.

By Matteo He, William F. Shen, Xinchi Qiu, Nicholas D. Lane
arXiv Computer Vision
Sep 21

From Retrieval to Recognition:How Vision--Language Models Become OCR Specialists

The paper investigates how general vision‑language models (VLMs) develop specialized optical character recognition (OCR) capabilities. By applying a causal intervention protocol, the authors identify sparse, stable OCR‑head sets in several VLMs and show that these heads largely overlap with textual retrieval/copy heads found in general VLMs. The study concludes that full‑sequence OCR functions as a dense multimodal copy‑and‑paste mechanism, and that when a VLM is fine‑tuned for OCR, it largely preserves the same head identities while redistributing their functional and causal strengths.

By Yuanxiang Huangfu, Hanmeng Zhong, Linqing Chen, Jeffrey Tiong Jee Hui
arXiv AI
4d ago

Similar Choices, Different Attention: Cross-Modal Associations in Humans and Vision-Language Models

The study compares human and vision‑language model (VLM) responses to cross‑modal association tasks, using identical stimuli (a pseudo‑word and two images) and recording both choices and eye movements. While larger VLMs show some alignment with human choices, their attention patterns correlate poorly with human gaze, performing no better than a simple center‑bias baseline. Fine‑tuning VLMs on human choices improves choice alignment but not attention alignment, and training on human gaze improves attention correlation without affecting choice accuracy.

By Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li