The paper investigates whether contrastive activation directions in large language models truly capture a functional valence axis or merely reflect correlated features of the contrast used to extract them. By conducting controlled interventions in a maze task across multiple LLM checkpoints, the authors show that directions fitted on one explicit outcome encoding transfer well to another, suggesting the readout is not tied to surface form. However, when the same outcome is achieved through announced versus unannounced histories, transfer degrades, indicating the direction remains strongly conditioned on prior announcements and does not represent a history‑invariant scalar valence state.
By Weihan Li, Xinlei Chen, Yuhan Song, Xiaofeng Lin, Tianshi Zheng
The paper introduces the Memory Decision Layer (MDL), a zero‑parameter controller that sits between retrieval and generation in large language models. MDL uses a three‑signal complementary encoder—combining relevance, reliability, and task risk—to produce an interpretable decision about the trustworthiness of retrieved memories. By decoupling confidence from consistency and enabling risk inversion and abstention, MDL cuts hallucination rates under conflicting memories by roughly 56% and nearly eliminates them in high‑risk scenarios, all while adding only 0.14 ms per decision.
By Yiming Zhang, Jinghong Zhang, Haoran Zhao, Yiren Ma, Chunlei Zhao
arXiv:2609.01048v1 Announce Type: cross
Abstract: Across the full Pythia suite (160M-12B, eight checkpoints, four task families), a linear probe can read a target variable from the residual stream as...
By Xining Xun
The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.
By Donna Vakalis
arXiv:2607. 06503v1 Announce Type: new Abstract: Large language model (LLM) agents solving multi-step tasks frequently commit to trajectories that are doomed to fail, yet continue to consume substantial inference compute before the failure becomes observable.
By Kai Ruan, Zihe Huang, Ziqi Zhou, Qianshan Wei, Xuan Wang, Hao Sun
arXiv:2607. 04419v3 Announce Type: replace Abstract: When evaluator-derived step rewards are pooled or compared across scoring channels, their sign is treated as transportable.
By Andrew Zhang, Chengzhan Li
arXiv:2606. 28525v1 Announce Type: cross Abstract: Fine-tuning on harmless data can partially undo behaviors acquired earlier in training.
By Samuele Poppi, Nils Lukas
arXiv:2608.30650v1 Announce Type: new
Abstract: LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn...
By Jie Liang, Zhengxin Yu, Hamid Nasiri, Peter Garraghan
arXiv:2608. 14717v1 Announce Type: cross Abstract: A query-relation deletion can improve the edited slot while reducing the utility of the prediction set that contains it.
By Ze Zhang, Yang Zhang
arXiv:2609.35808v1 Announce Type: cross
Abstract: Experience reuse can reduce repeated exploration in embodied agents, but a trajectory that succeeded previously may be unsuitable for the current exe...
By Quanquan Li, Hongbo Zhang, Yihe Chi, Liuyang Song, Jingyu Li, Yuxiang Huang, Hongzhen Zhang, Guitao Cao
arXiv:2608. 04788v1 Announce Type: cross Abstract: Large language model agents are commonly trained through reinforcement learning with sparse trajectory-level rewards, which offer limited guidance on how strongly individual tokens should be updated.
By Yi Yang, Cong Qin, Xiaodan Liu, Chishui Chen, Qing Dong, Yan Zhang, Cao Liu, Zhao Yang, Lu Pan, Jiaye Lin, Yi Feng
arXiv:2608.20442v1 Announce Type: new
Abstract: Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically...
By Qinyang Xu