arXiv:2607. 10441v1 Announce Type: cross Abstract: Context engineering decides what information a model carries forward, and current designs meter it in tokens: compressing the past into a bounded recurrent state, keeping a key-value entry for every token, or imposing a fixed budget through a window or eviction rule.
By Siddharth Pal, Viktoria Rojkova
arXiv:2603. 20381v2 Announce Type: replace-cross Abstract: Understanding the fundamental mechanisms governing the production of meaning in the processing of natural language is critical for designing safe, thoughtful, engaging, and empowering human-agent interactions.
By Christopher J. Agostino, Quan Le Thien, Nayan D'Souza, Louis van der Elst
Large Language Models (LLMs) are frequently portrayed as general-purpose solvers capable of solving arbitrary tasks. We argue that this view overlooks a fundamental constraint: language is a compressed and capacity-limited interface for conveying task information.
The paper introduces a pragmatic information theory that unifies communication, control, and decision-making through the isoteleia mapping, which formalizes equifinality by treating distinct semantic paths that lead to the same optimal action as pragmatically equivalent. It establishes a three-tier hierarchy of syntactic, semantic, and pragmatic information, defines pragmatic entropy, mutual information, channel capacity, and rate-distortion, and proves coding theorems that generalize Shannon’s results. The authors also present pragmatic value and cost of information, a Lagrangian dual framework for cross-layer optimization, and a pragmatic efficiency bound that quantifies the maximum net utility for resource-constrained intelligent systems, extending the theory to continuous messages and dynamic settings.
By Kai Niu, Ping Zhang
The paper investigates how language models decide between contextual information and their internal memory when the two conflict. By estimating "authority directions" from agreement prompts and swapping these directions between matched prompts, the authors show that such interventions can reproduce 30–68% of the shift in source choice across Qwen, Llama, and OLMo models. Cross‑task experiments reveal that authority directions learned on one task transfer only modestly (≈9%) to another, indicating that authority computations are largely task‑specific.
By Benjamin Shih, John Winnicki, Arianna Cao
arXiv:2607. 12216v1 Announce Type: cross Abstract: Multi-agent and memory-augmented LLM systems often place coordination content, shared state, prior discussion, tool outputs, summaries, and role instructions, inside the same finite prompt used for the current task.
By Brenda Lelis, Rodrigo Cabral-Carvalho
arXiv:2604. 09670v2 Announce Type: replace-cross Abstract: Intelligent systems must maintain and manipulate task-relevant information online to adapt to dynamic environments and changing goals.
By Hua-Dong Xiong, Li Ji-An, Jiaqi Huang, Robert C. Wilson, Kwonjoon Lee, Xue-Xin Wei
arXiv:2608. 15022v1 Announce Type: new Abstract: Language models hold latent quantities in a form they can report on, and more of a quantity is present in that form when the task requires reusing it flexibly.
By Parsa Mazaheri
arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.
By Kareem Hassani, Chaymaa Abbas, Lama Mawlawi, Mariette Awad
arXiv:2605. 18909v2 Announce Type: replace Abstract: Any system that models the world under finite representational capacity must compress; any compression entails a prior; and the prior is the system's bias.
By Ahmed Gamal Eldin
arXiv:2608. 01548v1 Announce Type: cross Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition.
By Yi Liu
The paper investigates how continuous latent states in large language models can store multiple reasoning steps through superposition. It challenges the intuition that retaining only the current reasoning frontier is optimal, showing that cumulative superposition of the full reasoning history can actually require fewer representational dimensions. The authors demonstrate that this approach preserves more valid evidence, improves downstream outcome discrimination, and delays unreliability, while also establishing that uniform cumulative weighting of memories is minimax‑optimal for future reasoning.
By Hongyu Gu, Chang Liu, Jingwen Fu