arXiv:2609.06036v1 Announce Type: new
Abstract: Proposal-based controllers---learned policies, language-model planners, and other black-box \emph{generators}---are increasingly deployed behind runtim...
By Guangxi Wan, Yongbo Xie, Yuqi Liu, Qingwei Dong, Qingxin Li, Hongfei Bai, Peng Zeng
arXiv:2607. 18305v1 Announce Type: cross Abstract: Some limits on what language models know are not gaps in data coverage but structural properties of learning from text.
By Priyansh Srivastava, Romit Chatterjee
The paper investigates whether a small, directly addressable change in the hidden state of a learned world model can steer its future predictions along a desired counterfactual trajectory. Using a 192‑dimensional recurrent model in a two‑object collision setting, the authors identify low‑rank latent carriers—specifically a rank‑4 patch—that, when applied, successfully redirect a 12‑step autonomous rollout without further intervention. The study demonstrates that this compact intervention interface consistently works across independently trained checkpoints and intervention times, while various control experiments confirm the specificity of the effect.
By Yang Liu, Yuming Chen
arXiv:2609.37841v1 Announce Type: new
Abstract: Masked generative models offer parallel token prediction, but accurate parallel sampling must account for dependencies among tokens. When dependencies...
By Ryotaro Kawata, Satoshi Hayakawa, Taiji Suzuki
arXiv:2606. 25449v1 Announce Type: cross Abstract: A language model's memory can be worse than having no memory at all.
By Alex Kwon
BARRIER (Bounded Activation Regions for Robust Information Erasure) is a method for machine unlearning that confines parameter updates to a controlled activation space, allowing stronger erasure of targeted concepts while limiting collateral damage to other representations. By employing interval arithmetic, it derives a closed‑form bound on worst‑case representation changes in protected regions, which serves as a knowledge‑preservation objective. The approach is architecture‑agnostic, compatible with existing erasure objectives, and empirically shows competitive performance in both classification and generative tasks, with enhanced robustness against adversarial recovery attacks.
By Jan Miksa, Patryk Krukowski, Przemys{\l}aw Spurek, Dawid Damian Rymarczyk, Marcin Sendera