From Exposure to Expectation: Frequency, Surprisal, and Language Across Development in Spanish
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The study investigates how disrupting conceptual versus referential information in short narratives affects human reading and large language model (LLM) processing. In humans, conceptual disruptions cause a strong, localized processing cost that peaks early and declines quickly, while referential disruptions produce weaker, gradually decreasing effects that are more influenced by sentence boundaries. In LLMs, both disruptions appear immediately at the manipulated word; surprisal patterns mirror human reading, whereas output-layer representations show that referential disruption initially causes a larger displacement before both types decay following a power-law.
arXiv:2505.12196v2 Announce Type: replace Abstract: The impressive linguistic abilities of large language models (LLMs) have recommended them as models of human sentence processing, with some conject...
Surprisal theory holds that the human processing difficulty of a linguistic unit in context is an affine function of its surprisal under some language model. I argue this claim is a tautology without further constraint: for any non-negative difficulty measure over units in context, there exists a language model whose surprisal is an affine function of it under mild technical conditions.
arXiv:2608. 14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior.
arXiv:2606. 07555v1 Announce Type: cross Abstract: Glossaries, technical specifications, and system prompts routinely ask language models to use familiar words in unfamiliar ways.
arXiv:2606. 11371v1 Announce Type: cross Abstract: Spoken language, whether produced by humans or large language models (LLM), unfolds over time with varying semantic content.