The paper investigates energy-based transformers as predictors of reading difficulty, extending the use of transformer language models in psycholinguistics. It demonstrates that the energy measure from these models robustly predicts reading times across multiple corpora, outperforming traditional metrics like surprisal and attention entropy. In a controlled experiment on relative clause processing, energy captures known asymmetries, suggesting it may unify previously complementary predictors.
By Jakub Dotlacil, Ece Takmaz
arXiv:2608. 14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior.
By Jinglei Ren, Yuyue Wang
The paper examines whether adding syntactic and rhetorical structure to text can improve the prediction of incoherence in large language model outputs. Experiments show that plain text actually yields higher accuracy, as the added structural information conflicts with the models’ architectures. The authors also demonstrate that coherence assessment can help detect misleading content by applying zero‑shot experiments to a Brazilian disinformation dataset.
By Victor Mazzotti, Luiz Pereira, Marina Bitencourt dos Santos, Helena Maia, Carlos Caetano, N\'adia Felix, Sandra Avila
arXiv:2608. 04021v1 Announce Type: cross Abstract: Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits.
By Han-yu Wang
The paper studies structural priming in language model production by conducting controlled sentence‑completion experiments on dative constructions. Results show that language models exhibit priming effects, especially when sentences are semantically coherent, with stronger relative increases for double‑object datives and larger absolute increases for prepositional‑object datives. The study also finds that primed completions involve more lexico‑semantic repetition, indicating that priming operates across syntactic, lexical, and semantic levels.
By Giulia Pucci, Ruizhe Li, Arabella Sinclair
arXiv:2511. 21338v2 Announce Type: replace Abstract: Masked Diffusion Language Models (MDLMs) have recently emerged as a promising alternative to Autoregressive Language Models (ARLMs), leveraging a denoising objective that, in principle, should enable more uniform context utilisation.
By Julianna Piskorz, Cristina Pinneri, Alvaro Correia, Motasem Alfarra, Risheek Garrepalli, Christos Louizos
arXiv:2609.18011v1 Announce Type: new
Abstract: In collaborative tasks with asymmetric information, participants coordinate their understanding through interaction. We ask whether gaze provides evide...
By Nan Li, Albert Gatt, Massimo Poesio
arXiv:2606. 07555v1 Announce Type: cross Abstract: Glossaries, technical specifications, and system prompts routinely ask language models to use familiar words in unfamiliar ways.
By Han-yu Wang
The study investigates how disrupting conceptual versus referential information in short narratives affects human reading and large language model (LLM) processing. In humans, conceptual disruptions cause a strong, localized processing cost that peaks early and declines quickly, while referential disruptions produce weaker, gradually decreasing effects that are more influenced by sentence boundaries. In LLMs, both disruptions appear immediately at the manipulated word; surprisal patterns mirror human reading, whereas output-layer representations show that referential disruption initially causes a larger displacement before both types decay following a power-law.
By Rui He, Nihal Altay, Wolfram Hinzen
arXiv:2603.04419v3 Announce Type: replace-cross
Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify...
By Murad Farzulla
arXiv:2606. 17389v1 Announce Type: cross Abstract: Multimodal Foundation Models are increasingly used as reasoning agents, making reliability, knowing when a model may hallucinate, critical.
By Logan Mann, Yi Xia, Ajit Saravanan, Ishan Dave, Saadullah Ismail, Shikhar Shiromani, Emily Huang, Ruizhe Li, Kevin Zhu
arXiv:2605. 10828v2 Announce Type: replace Abstract: As large language models are increasingly deployed in retrieval-augmented generation and agentic systems that accumulate extensive context, understanding how distracting information affects long-context performance becomes critical.
By Muhan Gao, Zih-Ching Chen, Kuan-Hao Huang