arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
The paper investigates where the ‘refusal’ behavior of language models resides across different architectures. It finds that a single direction in the residual stream governs refusal in transformers, and that the same direction—after a rigid rotation—also governs refusal in state‑space models (SSMs). By aligning these directions and applying a detector‑triggered gate, the authors demonstrate that refusal can be effectively transferred across transformer, SSM, recurrent, and hybrid architectures, showing that safety tooling can be ported by re‑estimating the direction at each architecture’s write site rather than rebuilding it from scratch.
By Preethi Carmel Bosco, Gopalakrishnan Srinivasan
arXiv:2607. 16895v1 Announce Type: new Abstract: Safe adaptive control is online adaptation under a safety guarantee on the learning trajectory itself.
By Venkatesh Saligrama
The paper demonstrates that safety mechanisms for autonomous large language model agents fail to compose across iterative loops, as trajectory‑scoped monitors cannot detect attacks whose evidence is spread over multiple iterations. It introduces LoopHarness, a system that maintains a persistent, non‑decaying safety state across loops, bounding unauthorized actions with a constant that does not grow with the number of iterations. The authors provide a comprehensive evaluation protocol, including attacks that require cross‑iteration evidence, module ablations, and adaptive white‑box red‑team testing.
By Chenhao Wu, Haoxuan Jia, Yang Liu, Yingguang Yang, Yuhan Lin, Chongyang Zhang, Hao Zheng, Yulin Huang, Jianshen Zhang, Yongzhi Qi, Shang Luo, Kefu Xu, Jifeng Zhu, Bin Chong
arXiv:2608. 12791v1 Announce Type: cross Abstract: What a finite learning device has recorded and what will hold value for it on future tasks are not the same quantity.
By Akihito Sudo
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty.
arXiv:2608.27782v1 Announce Type: cross
Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is t...
By Xujun Che, Depeng Xu, Shuhan Yuan
arXiv:2608. 19491v1 Announce Type: new Abstract: Most modern optimizers form their momentum as an exponential moving average (EMA) of past gradients, forgetting every direction at one fixed rate.
By Euijin Hong, Guannan Qu
arXiv:2609. 11326v2 Announce Type: replace-cross Abstract: We ask whether it can be certified algorithmically that a self-modifying computational system preserves a safety property at its next step (preservation) and along its whole evolution (persistence).
By Jose Pascual Gumbau Mezquita
arXiv:2608. 19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored.
By Zhiqiang Tan
arXiv:2608. 26515v1 Announce Type: cross Abstract: We study online prediction for a specific finite-alphabet, exogenously driven source with infinite input memory.
By Vaneet Aggarwal
arXiv:2606. 21253v2 Announce Type: replace Abstract: Continual learning that is gradient-free, local, online, and append-only is attractive for edge and streaming deployment, but its value is usually argued informally.
By Jianwei Lou (RailMind Systems, Neuss, Germany)