A Systematic Analysis of the Predictive Power of LM Surprisal in Reading Chinese
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The study examines how reader proficiency influences the relationship between layer-wise surprisal from large language models (LLMs) and eye-tracking gaze measures. Using the MECO L2 corpus, researchers compared high- and low-proficiency readers on first-pass gaze duration (FPGD) and total gaze duration (TGD), finding that lower-proficiency readers exhibit deeper Predictive Depth for FPGD, while TGD shows deeper Predictive Depth across both groups. The results suggest that the distribution of predictive power across LLM layers relates to the timing and breadth of reading processes and varies with reader proficiency.
The study investigates whether language models trained specifically on Cantonese better predict human reading patterns by comparing eye-tracking data with information-theoretic metrics derived from several models. Two adaptation contrasts were examined: a lightly adapted CKIP GPT-2 Tiny versus its Cantonese derivative JED351, and a heavily adapted Qwen2.5-7B versus CantoneseLLM-7B. Results show that the extensively Cantonese-trained CantoneseLLM-7B consistently outperforms others on lexical surprisal and joint metrics, while entropy reduction favors the less adapted CKIP model, indicating that deeper Cantonese-specific training can improve predictive alignment but that rankings vary by metric.
arXiv:2505.12196v2 Announce Type: replace Abstract: The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conject...
arXiv:2607. 08152v1 Announce Type: cross Abstract: On the recent EyeBench benchmark, predicting reading comprehension from eye movements exposes a stark gap: text-aware models using pretrained language models reach 56--63% AUROC, while gaze-only models operate at chance.
arXiv:2610.08879v1 Announce Type: new Abstract: Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Mask...
arXiv:2607. 03994v1 Announce Type: cross Abstract: Modern language models generally represent text as sequences of discrete token embeddings, an assumption deeply rooted in current practice but rarely questioned.