arXiv:2609.36081v1 Announce Type: new
Abstract: Representations continually change as a network learns new tasks. We ask whether early representational changes naturally form a geometric structure th...
By Yuantao Deng, Jinnuo Liu, Kaizhen Tan, Yuchen Liu
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
By Akarsh Kumar, Phillip Isola
The paper investigates continual machine unlearning, where models must forget data over time. It identifies a fundamental issue called plasticity collapse, where successive unlearning requests cause geometric constraints that saturate parameter space, leading to two failure modes: forward failure (reduced forgetting quality) and backward failure (re‑memorization). Experiments across architectures and datasets confirm that plasticity collapse is a pervasive problem in continual unlearning.
By Yingdan Shi, Xiang Xu, Kaize Ding, Alfred O. Hero, Ren Wang
arXiv:2606. 20431v1 Announce Type: new Abstract: Continual learning (CL) systems often forget previously acquired knowledge, yet the mechanisms driving forgetting remain hard to isolate in practice because real datasets entangle many factors.
By Jan Wasilewski, J\k{e}drzej Kozal, Micha{\l} Wo\'zniak, Bartosz Krawczyk
arXiv:2608. 11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting.
By Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
arXiv:2607. 11958v1 Announce Type: new Abstract: Under the free energy principle, a predictive system does not observe reality directly; it maintains a generative model of the world and experiences that model's best current hypothesis.
By MD Ibrahim Hossain Ridoy
The paper presents a non-equilibrium dynamical mean-field theory (DMFT) that explains how learning reshapes the dynamics of recurrent neural networks, turning initially chaotic activity into stable, task-dependent behavior. It shows that a slow, feedback-driven learning process gradually increases effective feedback strength, driving the network through a bifurcation that marks the transition from chaotic to stable dynamics. By deriving the two-time correlation function, the authors identify a critical feedback strength and a learning-rate-dependent critical time that separate these regimes, and they demonstrate that the theory accurately predicts the network’s output evolution during training, matching numerical simulations.
By Varun Vaidya
arXiv:2609.06006v1 Announce Type: cross
Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state...
By Minh Hoang Nguyen, Huu Hiep Nguyen, Manh Nguyen, Van Dai Do, Dung Nguyen, Hung Le
arXiv:2602. 18131v2 Announce Type: replace Abstract: Temporal Predictive Coding provides a layer-local, parallelisable mechanism for learning in recurrent systems, making it an attractive candidate for online local learning on neuromorphic and edge hardware.
By Tom Potter, Oliver Rhodes
arXiv:2607. 25531v1 Announce Type: cross Abstract: Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure.
By Zeki Doruk Erden
arXiv:2608. 04358v1 Announce Type: new Abstract: Continual learning (CL) requires models to learn tasks sequentially, yet deep neural networks often suffer from plasticity loss and poor knowledge transfer, which can impede their long-term adaptability.
By Seyed Roozbeh Razavi Rohani, Khashayar Khajavi, Wesley Chung, Mandana Samiei, Mo Chen
We introduce Neural Subspace Reallocation (NSR), which reframes continual learning as memory management over parameter subspaces. Instead of treating Low-Rank Adaptation (LoRA) modules as disposable per-task adapters, NSR manages them as compressible, retrievable memory units on a frozen backbone through a recurring cycle: (1) compress learned LoRAs via SVD, (2) reserve them in a TaskKnowledgeBank, (3) recall related past LoRAs by embedding similarity to warm-start new or returning tasks, and (4) reallocate the active subspace accordingly, with distillation protecting prior tasks.