arXiv Machine Learning By Linzhe Zhang, Changming Xu

What Can a Recurrent State Safely Forget?

Read the original on arXiv Machine Learning →

arXiv:2609. 23366v1 Announce Type: new Abstract: Recurrent models must preserve information that changes future behavior while suppressing hidden-state error.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 18

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

The paper investigates whether a small, directly addressable change in the hidden state of a learned world model can steer its future predictions along a desired counterfactual trajectory. Using a 192‑dimensional recurrent model in a two‑object collision setting, the authors identify low‑rank latent carriers—specifically a rank‑4 patch—that, when applied, successfully redirect a 12‑step autonomous rollout without further intervention. The study demonstrates that this compact intervention interface consistently works across independently trained checkpoints and intervention times, while various control experiments confirm the specificity of the effect.

By Yang Liu, Yuming Chen
arXiv Computer Vision
Sep 25

BARRIER: Bounded Activation Regions for Robust Information Erasure

BARRIER (Bounded Activation Regions for Robust Information Erasure) is a method for machine unlearning that confines parameter updates to a controlled activation space, allowing stronger erasure of targeted concepts while limiting collateral damage to other representations. By employing interval arithmetic, it derives a closed‑form bound on worst‑case representation changes in protected regions, which serves as a knowledge‑preservation objective. The approach is architecture‑agnostic, compatible with existing erasure objectives, and empirically shows competitive performance in both classification and generative tasks, with enhanced robustness against adversarial recovery attacks.

By Jan Miksa, Patryk Krukowski, Przemys{\l}aw Spurek, Dawid Damian Rymarczyk, Marcin Sendera