arXiv AI

Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans

arXiv:2608. 14681v1 Announce Type: cross Abstract: Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior.

arXiv AI
Sep 7

Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation

The paper studies structural priming in language model production by conducting controlled sentence‑completion experiments on dative constructions. Results show that language models exhibit priming effects, especially when sentences are semantically coherent, with stronger relative increases for double‑object datives and larger absolute increases for prepositional‑object datives. The study also finds that primed completions involve more lexico‑semantic repetition, indicating that priming operates across syntactic, lexical, and semantic levels.

By Giulia Pucci, Ruizhe Li, Arabella Sinclair
arXiv AI
Sep 1

Mechanism Shift During Post-training from Autoregressive to Masked Diffusion Language Models

The study investigates how post‑training of large autoregressive language models (ARMs) into masked diffusion models (MDMs) affects their internal computation. Across two 7B ARM‑MDM families and four diagnostic tasks, the authors find that MDMs retain much of the ARM’s high‑attribution pathways on prefix‑dominant tasks, but reorganize computation toward earlier layers on globally constrained tasks. Component‑level probes reveal that ARMs depend on sharply specialized components, whereas MDMs show weaker specialization and more diffuse output‑space alignment.

By Injin Kong, Hyoungjoon Lee, Yohan Jo
arXiv AI
Sep 25

Post-Training Leaves Behavioral Shadows on Unrelated Decisions

The paper demonstrates that language models can acquire new capabilities from post‑training data even when the training text is unrelated to the target task. Using a method called Active Taskless Distillation (ATD), the authors show that a single word from a teacher model can transfer knowledge to a student model without any target‑task examples or teacher logits. Experiments on Qwen2.5-1.5B reveal significant performance gains on HumanEval+ and improvements in scientific knowledge, commonsense reasoning, and reading comprehension across various model families.

By Ziyang Zhang, Yubin Jing, Yuanhao Zeng, Yuyao Li, Haofan Wang, Yichen Gong
arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
arXiv AI
Aug 14

From Observation to Intervention: Memory in Brains and Large Language Models

arXiv:2608. 12377v1 Announce Type: cross Abstract: Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader associations, how new information is written or updated, and how memory-related states can be perturbed.

By Morteza Salehjahromi, Shayan A. Zadegan, Amgad Muneer, Jia Wu
arXiv AI
Sep 3

Language Models Can Control Their Own Attention

The paper introduces Declarative Attention (DA), a protocol that lets language models explicitly declare which parts of their context to focus on during generation. By partitioning decoding into full-context, region-specific, and recent-output-only modes, the inference engine can skip large portions of the KV cache, dramatically reducing attended tokens. Experiments on 15 long-context tasks with off-the-shelf models show significant savings (52.0% and 31.1% reductions) with only modest accuracy drops that diminish as model size increases.

By Namgyu Ho, Huzama Ahmad, Woosung Koh, Se-Young Yun, Tal Schuster, Cicero Nogueira dos Santos