arXiv Machine Learning

Early Learning Shapes Later Directions Of Representation Change In Continual Learning

arXiv AI
Jun 17

Dimensionality Controls When Modularity Helps in Continual Learning

arXiv:2606. 17889v1 Announce Type: cross Abstract: Compositional learning systems must balance plasticity, the ability to acquire new knowledge, with stability, the preservation of previously learned components, especially when tasks share structure and risk interference.

By Kathrin Korte, Christian Medeiros Adriano, Joachim Winther Pedersen, Eleni Nisioti, Sebastian Risi
arXiv Computer Vision
Aug 27

A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training

The paper introduces a new task called Multimodal Unsupervised Continual Post-Training (MU‑CPT), which allows multimodal large language models (MLLMs) to continuously learn from streaming unlabeled data. It identifies token‑level visual dependence (VD) as essential for MU‑CPT, using its structural distortion to detect cross‑modal forgetting and its heterogeneity to guide new‑task learning. The proposed Visual Dependence‑Aware (VDA) framework includes Visually Constrained Optimal Transport (VC‑OT) to mitigate forgetting and Visually Modulated Adaptation (VMA) to enhance new‑task plasticity, achieving a balance between stability and adaptability in MU‑CPT.

By Kaichen Li, Zhilin Zhu, Jianhao Huang, Zhengqin Lai, Baochen Xiong, Zibo Shao, Yaguang Song, Linhui Xiao, Xiaoshan Yang, Changsheng Xu
arXiv Machine Learning
Sep 2

How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks

The paper analyzes how temporal correlations in input sequences affect memory formation in linear recurrent neural networks (LRNNs). By solving the learning dynamics for correlated inputs, it shows that correlation introduces a cost to retaining past information, reshaping the learning trajectory and reducing the network’s memory of past inputs. Key findings include a correlation‑dependent threshold for memory retention, the influence of input similarity on memory usefulness, and the emergence of a feedthrough path when zero error is required.

By Arnol Manuel Fokam, Fasseu Sieyondji Akpevwoghene, Edem Fiifi Dawson
arXiv AI
Sep 18

Why $\beta_1 = \beta_2$ Is Dynamically Special in Adam

The paper investigates why setting the two momentum parameters of Adam equal (β1=β2) has a special dynamic effect. By analysing Adam in continuous time, the authors show that the update decomposes into a sign component, a magnitude‑lag term proportional to the difference between the two memory times, and other terms. This lag term disappears exactly when β1=β2, making the diagonal the only regime where the mismatch‑induced response is structurally absent. Experiments on six vision and language tasks confirm that tied configurations are sign‑dominated, have smaller lag contributions, and exhibit smoother update‑norm trajectories.

By Alberto Fern\'andez-Hern\'andez, Cristian P\'erez-Corral, Jose I. Mestre, Manuel F. Dolz, Enrique S. Quintana-Ort\'i