The paper investigates what a steerable valence direction in large language models (LLMs) actually represents, focusing on a good‑bad outcome direction in a maze task. By using controlled interventions that separate the realized outcome from the informational history that led to it, the authors find that directions trained on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer performance drops sharply, indicating that the post‑event readout remains strongly conditioned on the earlier announcement. In a matched maze‑reinforcement‑learning run, the post‑RL direction becomes more predictive of reference‑MDP return and the policy depends more on it, yet the history dependence persists. These findings support a functional, value‑related interpretation of the direction but argue against identifying it with a history‑invariant scalar valence state.
By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
arXiv:2605. 26772v1 Announce Type: cross Abstract: Large reasoning models (LRMs) generate chain-of-thought (CoT) traces before producing final outputs, introducing a dynamic internal state that may complicate control mechanisms such as refusal.
By Kia-J\"ung Yang, Dominik Meier, Jiachen Zhao, Terry Ruas, Bela Gipp
The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.
By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao
arXiv:2607. 04419v3 Announce Type: replace Abstract: When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable.
By Andrew Zhang, Chengzhan Li
arXiv:2607. 10608v1 Announce Type: new Abstract: Memory is becoming a core component of long-horizon AI agents, allowing agents to reuse past experience when operating web browsers, software tools, and other interactive environments.
By Yixiong Chen, Xinyi Bai, Alan Yuille
arXiv:2608.30650v1 Announce Type: new
Abstract: LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn...
By Jie Liang, Zhengxin Yu, Hamid Nasiri, Peter Garraghan
arXiv:2605.12978v2 Announce Type: replace
Abstract: Learning from past experience benefits from two complementary forms of memory: episodic traces -- raw trajectories of what happened -- and consolid...
By Dylan Zhang, Yanshan Lin, Zhengkun Wu, Yihang Sun, Bingxuan Li, Dianqi Li, Hao Peng
arXiv:2608. 06811v1 Announce Type: cross Abstract: Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration, hypothesis, implementation, and verification.
By Jiahao Zhang, Yifan Zhang, Yu Huang
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
arXiv:2609.01048v1 Announce Type: cross
Abstract: Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as...
By Xining Xun
arXiv:2607. 18114v1 Announce Type: cross Abstract: Modern LLMs are alarmingly susceptible to surprisingly simple immaterial changes of input prompts: a casual hint, an incorrectly labeled few-shot example, or a fake prior assistant turn often flips an originally correct answer.
By Prakhar Gupta, Terry Jingchen Zhang, Florent Draye, Bernhard Sch\"olkopf, Zhijing Jin