arXiv:2601.12075v2 Announce Type: replace
Abstract: Language models used in retrieval-augmented settings must arbitrate between parametric knowledge stored in their weights and contextual information...
By Mehrdad Farahani, Franziska Penzkofer, Richard Johansson
arXiv:2607. 09889v1 Announce Type: cross Abstract: Fixed-state sequence models compress an unbounded past into a bounded state, which caps their associative recall at roughly the state dimension; attention escapes the cap by keeping a key-value entry for every token, at quadratic compute and a cache that grows with the sequence.
By Siddharth Pal, Viktoria Rojkova
arXiv:2606. 26806v1 Announce Type: new Abstract: Long-running language agents need more than memory access.
By Haoliang Han
arXiv:2603. 23530v2 Announce Type: replace-cross Abstract: Large language models often fail to satisfy formatting instructions when they must simultaneously perform demanding tasks.
By Avni Mittal
The paper introduces an action‑class diagnostic framework for multi‑turn tool‑calling in large language model agents, breaking failures into action‑class miscalibration and action‑execution failure across a four‑class action space (TOOL_CALL, ASK, REFUSE, CONFIRM). It defines a self‑revealing upper bound (Acc GAR) to expose state‑grader masking of miscalibration and shows that miscalibration is a significant, previously hidden failure mode, especially for heavily tool‑trained families. The study demonstrates that calibration can be reshaped by context‑only perturbations, but the effects vary widely across models and perturbation mechanisms, underscoring the need for diagnostics beyond aggregate accuracy.
By Kangjia Zhao, Jiajun Li, Haozhan Shen, Wei Chow, Linfeng Li, Hang Song, Lingdong Kong, Chen Zhi, Tiancheng Zhao, Songhua Liu, Jianwei Yin
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
arXiv:2607. 17545v1 Announce Type: new Abstract: Language agents depend on memory across interactions.
By Qingcan Kang, Mingyang Liu, Shixiong Kai, Kaichao Liang, Zhentao Tang, Yuqi Cui, Tao Zhong, Mingxuan Yuan
The paper investigates in-context binding errors in language models, showing that a linear probe can recover correct entity bindings from frozen hidden states even when the model outputs incorrect bindings. Across 16 checkpoints, the probe’s accuracy on failure cases surpasses a baseline by about 0.196, and a probe‑based score improves failure detection over the model’s confidence by 0.079 AUROC. Steering the residual stream toward the probe‑decoded binding further boosts accuracy by an average of 0.168 across eight models.
By Manas Venkata Sai Ravulapalli, Samrath Singh Chadha, Abhinav M. Hari
arXiv:2607. 14112v1 Announce Type: cross Abstract: Large language models (LLMs) are evaluated as though perfect reliability is achievable for any task given sufficient scale.
By Subhabrata Majumdar
arXiv:2609.37076v1 Announce Type: new
Abstract: Large language models trained on vast corpora inherently risk memorizing harmful content that may later re-emerge in their outputs. To mitigate this is...
By Puning Yang, Qizhou Wang, Junchi Yu, Bo Han, Xiuying Chen
arXiv:2507. 01414v2 Announce Type: replace Abstract: We introduce a new family of toy problems that combine features of linear-regression-style continuous in-context learning (ICL) with discrete associative recall.
By Sultan Daniels, Dylan Davis, Dhruv Gautam, Wentinn Liao, Gireeja Ranade, Anant Sahai
arXiv:2607. 08032v1 Announce Type: new Abstract: Large language models, and the agents built on them, spend an ever-growing share of their compute and memory on remembering: caching attention keys and values, carrying long prompts, maintaining recurrent state, and storing what happened in previous turns and sessions.
By Ashwin Gerard Colaco, Nada Lahjouji