arXiv:2607. 25531v1 Announce Type: cross Abstract: Contemporary machine learning struggles to learn continually, reuse prior knowledge, and expose a comprehensible internal structure.
By Zeki Doruk Erden
arXiv:2606. 17889v1 Announce Type: cross Abstract: Compositional learning systems must balance plasticity, the ability to acquire new knowledge, with stability, the preservation of previously learned components, especially when tasks share structure and risk interference.
By Kathrin Korte, Christian Medeiros Adriano, Joachim Winther Pedersen, Eleni Nisioti, Sebastian Risi
arXiv:2606. 20431v1 Announce Type: new Abstract: Continual learning (CL) systems often forget previously acquired knowledge, yet the mechanisms driving forgetting remain hard to isolate in practice because real datasets entangle many factors.
By Jan Wasilewski, J\k{e}drzej Kozal, Micha{\l} Wo\'zniak, Bartosz Krawczyk
arXiv:2608. 15854v1 Announce Type: new Abstract: Catastrophic forgetting remains a fundamental obstacle to continual learning, where neural networks lose previously acquired knowledge while learning new tasks.
By Maksim A. Kazanskii
The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.
By Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu
arXiv:2609.37836v1 Announce Type: new
Abstract: Neural networks trained toward the same final objective can reach similar predictive performance while retaining internal representations shaped by ear...
By Ertu\u{g}rul Mutlu
The paper analyzes how temporal correlations in input sequences affect memory formation in linear recurrent neural networks (LRNNs). By solving the learning dynamics for correlated inputs, it shows that correlation introduces a cost to retaining past information, reshaping the learning trajectory and reducing the network’s memory of past inputs. Key findings include a correlation‑dependent threshold for memory retention, the influence of input similarity on memory usefulness, and the emergence of a feedthrough path when zero error is required.
By Arnol Manuel Fokam, Fasseu Sieyondji Akpevwoghene, Edem Fiifi Dawson
In this paper, we explore a novel task of Multimodal Unsupervised Continual Post-Training (MU-CPT), enabling deployed MLLMs to continually evolve from streaming unlabeled data. Existing unsupervised p...
arXiv:2602. 03846v2 Announce Type: replace-cross Abstract: We develop a continual learning method for pretrained models that \emph{requires no access to old-task data}, addressing a practical barrier in foundation model adaptation where pretraining distributions are often unavailable.
By Romain Cosentino
arXiv:2606. 06032v1 Announce Type: new Abstract: Catastrophic forgetting is commonly interpreted as the irreversible erasure of previously acquired knowledge during sequential learning.
By Ayushman Trivedi, Bhavika Melwani
arXiv:2608. 11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting.
By Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
The paper investigates why setting the two momentum parameters of Adam equal (β1=β2) has a special dynamic effect. By analysing Adam in continuous time, the authors show that the update decomposes into a sign component, a magnitude‑lag term proportional to the difference between the two memory times, and other terms. This lag term disappears exactly when β1=β2, making the diagonal the only regime where the mismatch‑induced response is structurally absent. Experiments on six vision and language tasks confirm that tied configurations are sign‑dominated, have smaller lag contributions, and exhibit smoother update‑norm trajectories.
By Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ort\'i